system

The system automates book summary generation and audio delivery, addressing the need for efficient content understanding in busy lives by providing summarized book content in audio format with personalized recommendations.

JP2026064592APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Current technologies lack a system that can automatically generate summaries of books and provide them in audio format, making it difficult for busy individuals to efficiently understand book content.

Method used

A system that allows users to input book information, search for the book, summarize the text, convert it into an audio file, and deliver it to their terminal, with options for playback and favorites list management, enabling efficient content acquisition.

Benefits of technology

Enables users to quickly and conveniently understand book content through automated summary generation and audio provision, facilitating easy access and personalized recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064592000001_ABST
    Figure 2026064592000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for users to input information about the books they want to read, A means for receiving information about a book entered by the user and searching for the book in a database, A means of summarizing the full text of the searched books, A means of converting the summarized content of a book into an audio file, A means for distributing the aforementioned audio file to the user's terminal, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] ---

[0005] In modern society, many people lead busy lives and it is difficult to secure time to read books they want to read. Also, there are many situations where they cannot read long texts, or they want to quickly acquire specific knowledge. Current technologies lack a system that can automatically generate a summary of a book and provide it in audio, and thus cannot meet such needs. The object of this invention is to provide a means for efficiently understanding the content of a book even in a busy life by integrating automatic generation of a summary and audio provision.

Means for Solving the Problems

[0006] The present invention provides a system that includes means for a user to input information about a book they wish to read, means for receiving the book information entered by the user and searching for the book in a database, means for summarizing the full text of the searched book, means for converting the summarized book content into an audio file, and means for delivering the audio file to the user's terminal. Furthermore, the system includes means for the user to select an option to play the audio file, enabling the summarized book content to be provided instantly in audio format. In addition, the system includes means for the user to register books they wish to read in a favorites list and to send notifications when new summaries are generated, allowing for efficient information acquisition. In this way, users can efficiently understand the content of books even in their busy lives.

[0007] ---

[0008] A "user" refers to an individual who uses the system to search for books and obtain a summarized version of the content in audio format.

[0009] "Book information" refers to identifying information used to identify a specific book, such as the title, author's name, and ISBN code.

[0010] A "database" refers to an electronic recording device that stores information such as the full text and metadata of books.

[0011] A "summary" refers to a short, concise text created by extracting the most important parts from the full text of a book.

[0012] An "audio file" refers to audio data generated using speech synthesis technology based on summarized content.

[0013] A "terminal" refers to an electronic device used by users to input information from books and play audio files.

[0014] "Distribution" refers to the act of sending audio files from a server to a terminal.

[0015] "The means for selecting a playback option" refers to an interface for a user to select a command for playing an audio file on a terminal.

[0016] "The means for registering in a favorite list" refers to a function for a user to temporarily save books of interest and create a list for later access.

[0017] "Notification" refers to the act of sending a message to inform a user when a new summary is generated.

Brief Description of Drawings

[0018] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of a data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Modes for Carrying Out the Invention

[0019] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0020] First, the language used in the following description will be explained.

[0021] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0022] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0026] [First Embodiment]

[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0039] ---

[0040] This invention provides a system that enables users to efficiently understand the content of books even in their busy lives. This system has the function of automatically generating a book summary and providing it in audio format. The program processing of this system is described in detail below.

[0041] System Overview

[0042] This system consists of a server and terminals, allowing users to input book information via the terminals and receive summarized content in audio format. The system has the following main functions:

[0043] 1. Book Selection

[0044] User: Enter the title or ISBN code of the book you want to read using the terminal's interface. For example, if the user wants to read "a famous novel," they would enter the title in the terminal's search bar.

[0045] 2. Searching for book data

[0046] Terminal: Sends the book information entered by the user to the server.

[0047] Server: Based on the received information, it searches the database and finds data for matching books. It then sends the found data back to the terminal. For example, it might send the full text and related metadata of a famous novel to the terminal.

[0048] 3. Automatic summary generation

[0049] Server: The full text of a book is passed to a natural language processing (NLP) engine, which extracts the important parts and generates a short summary. For example, the NLP engine extracts the key plot points and characters from the full text of a famous novel and summarizes them concisely.

[0050] 4. Generating audio files

[0051] Server: Inputs the generated summary into the speech synthesis engine and converts it into audio data. For example, the speech synthesis engine converts the summary content into an MP3 audio file.

[0052] 5. Audio distribution

[0053] Server: Saves the audio file to temporary storage and generates a download link.

[0054] Device: Presents the received download link to the user. For example, a notification might appear stating, "You can listen to an audio summary of a famous novel."

[0055] 6. Audio Playback

[0056] User: Select an audio playback option from the device's notifications or interface. For example, the user presses the "Play" button.

[0057] Device: Download the audio file from the download link and play it using the built-in player.

[0058] 7. Add to favorites and receive notifications

[0059] User: Adds books they want to read to their favorites list. For example, the user presses the "Add to Favorites" button.

[0060] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[0061] In this way, the system of the present invention provides users with a means to efficiently understand the content of books even in their busy lives. By automating each processing step, users can easily access and use the system. Although specific examples have been provided, the system is not limited to these and can be applied in a variety of ways.

[0062] The following describes the processing flow.

[0063] ---

[0064] Step 1:

[0065] User: Open the dedicated application or webpage on your device and enter the title or ISBN code of the book you want to read.

[0066] Step 2:

[0067] Terminal: Sends the book information entered by the user to the server.

[0068] Step 3:

[0069] Server: Based on the title or ISBN code of the received book, the server searches the database for data on matching books.

[0070] Step 4:

[0071] Server: Returns the full text of the searched book and related metadata to the terminal.

[0072] Step 5:

[0073] Server: Passes the full text to a natural language processing (NLP) engine, which extracts the important parts and generates a summary.

[0074] Step 6:

[0075] Server: Inputs the generated summary into the speech synthesis engine and converts it into audio data. Specifically, the speech synthesis engine converts the summary content into an MP3 audio file.

[0076] Step 7:

[0077] Server: Saves the generated audio files to temporary storage and generates a download link.

[0078] Step 8:

[0079] Server: Sends a download link to the device.

[0080] Step 9:

[0081] Terminal: Receives the sent download link and displays a notification to the user stating, "The summary is available for audio playback."

[0082] Step 10:

[0083] User: Select the audio playback option in the device's notification or interface and press the "Play" button.

[0084] Step 11:

[0085] Device: Download the audio file from the download link and play the audio using the built-in player.

[0086] Step 12:

[0087] User: Add books you want to read to your favorites list as needed. For example, the user presses the "Add to Favorites" button.

[0088] Step 13:

[0089] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[0090] ---

[0091] The above is a detailed explanation of the processing steps. This system automates the entire process, from inputting book information to audio playback, to help users efficiently understand the content of books.

[0092] (Example 1)

[0093] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] In today's busy lives, many people don't have the time to efficiently understand the content of books. There is a growing demand for systems that allow users to easily access summarized book content and listen to it as audio. Traditional systems lacked convenience because the processes of generating summaries, converting to audio files, and distribution were time-consuming and laborious.

[0095] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0096] In this invention, the server includes means for inputting information about a book that the user wants to read, means for receiving the book information input by the user and searching for the book in a data store, means for summarizing the entire text of the searched book, means for inputting the generated summary into a natural language processing engine, means for inputting the generated summary into a speech synthesis engine and converting it into audio data, means for temporarily storing the audio data and generating a download link thereof, and means for delivering the download link to the user's terminal. This makes it possible for the user to efficiently and quickly obtain an audio summary of the book.

[0097] A "user" is an individual or organization that uses the system to obtain book summaries and audio data.

[0098] "Book information" refers to information necessary to uniquely identify a book, such as the book's title, author's name, and ISBN code.

[0099] A "data store" is a storage device or database system used to store and manage book data.

[0100] "The full text of the book" refers to all the text and data contained in the book.

[0101] A "summary" refers to a concise compilation of the most important parts and key points extracted from the entire text of a book.

[0102] A "natural language processing engine" is software or algorithms that analyze text data and perform tasks such as summarizing, translating, and other natural language processing.

[0103] A "speech synthesis engine" is a software or hardware system that takes text data as input and converts it into speech data.

[0104] "Audio data" refers to audio files in an audible format, generated by a speech synthesis engine.

[0105] "Temporary storage" refers to the process of saving data only for a certain period, with deletion or updating performed as needed.

[0106] A "download link" refers to a URL or path used to obtain audio data or other data.

[0107] A "terminal" is an electronic device used by a user to access and operate a system.

[0108] This invention provides a system that enables users to efficiently understand the content of books even in their busy lives. This system has the function of automatically generating a book summary and providing it in audio format. Specific embodiments of this system are described in detail below.

[0109] System Configuration

[0110] This system primarily consists of a server and terminals, and users obtain book summaries as audio via the terminals. The system is comprised of the following hardware and software combinations.

[0111] hardware

[0112] The server includes a data store (e.g., a MySQL® database), an NLP engine, and a speech synthesis engine.

[0113] Device: An electronic device used by a user, such as a smartphone, tablet, or personal computer.

[0114] software

[0115] Natural language processing engine (e.g., GPT-3 (registered trademark))

[0116] Speech synthesis engine (e.g., Google® Text-to-Speech)

[0117] Database management system (e.g., MySQL)

[0118] HTTP request processing libraries (e.g., Axios)

[0119] Operation flow

[0120] The operation flow of this system is as follows:

[0121] 1. Book Selection

[0122] The user enters the title or ISBN code of the book they want to read through the terminal's interface. For example, if the user wants to read "a famous novel," they enter the title in the terminal's search bar and press the "Search" button.

[0123] 2. Searching for book data

[0124] The terminal sends the book information entered by the user to the server. Based on the received information, the server searches its data store for matching book data. As a result, it sends back the full text of "a famous novel" and related metadata to the terminal.

[0125] 3. Automatic summary generation

[0126] The server passes the full text of the book to a natural language processing engine, which extracts the important parts and generates a short summary. For example, it passes a prompt sentence like the following to the natural language processing engine:

[0127] Please summarize the following text: "The complete text of a famous novel."

[0128] 4. Generating audio files

[0129] The server inputs the generated summary into the speech synthesis engine and converts it into speech data. Specifically, it uses prompts like the following:

[0130] "Please convert this summary into an audio file."

[0131] The generated audio file will be temporarily saved.

[0132] 5. Audio distribution

[0133] The server generates a download link for the audio file and notifies the terminal of that link. For example, it sends the user a message saying, "You can listen to an audio summary of a famous novel." The terminal then displays the received download link on its interface.

[0134] 6. Audio Playback

[0135] The user selects the audio playback option through the device's notifications or interface. When the user presses the "Play" button, the device retrieves the audio file from the download link and plays it using the built-in player.

[0136] 7. Add to favorites and receive notifications

[0137] Users can add books they want to read to their favorites list. When a new summary is generated, the server notifies the user via email or app notification.

[0138] This system enables users to efficiently understand the content of books even in their busy lives, and by automating each processing step, it achieves simple access and use. For example, it can be used in the same process when a user is trying to read the latest business book, "The Road to Success."

[0139] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0140] Step 1:

[0141] User: Enter information about the book you want to read. Specifically, the user enters the book title and ISBN code through the terminal interface and presses the "Search" button. This entered information is then sent to the next processing step.

[0142] Step 2:

[0143] Terminal: Sends the book information entered by the user to the server. Specifically, it sends the book information to the server using an HTTP request. This operation passes the input data to the server.

[0144] Step 3:

[0145] Server: Searches the data store based on the received book information. Specifically, it queries the data store (e.g., a MySQL database) based on the received book information and retrieves the full text and metadata of matching books. This query process outputs the full text and metadata.

[0146] Step 4:

[0147] Server: The server passes the full text obtained as search results to a natural language processing (NLP) engine to generate a summary. Specifically, it inputs the full text into the NLP engine using the prompt message, "Summarize the following text: The full text of 'Book Title'." The NLP engine performs text analysis, extracts the important parts, and generates a summary. This summary is then passed to the next step.

[0148] Step 5:

[0149] Server: Inputs the generated summary into a text-to-speech engine and converts it into audio data. Specifically, it inputs the summary into a text-to-speech engine (e.g., Google Text-to-Speech) using a prompt message such as "Please convert this summary into an audio file." The text-to-speech engine converts the text into speech and generates an audio file in MP3 format. This audio file is then passed on to the next step.

[0150] Step 6:

[0151] Server: Temporarily stores the audio file and generates a download link. Specifically, it uploads the generated audio file to cloud storage and generates a download link. This link is then sent to the next step.

[0152] Step 7:

[0153] Device: Notifies the user of the generated download link. Specifically, the device's interface and notification system display the link along with the message, "You can listen to an audio summary of the book title." When the user clicks this link, they proceed to the next step.

[0154] Step 8:

[0155] User: Selects the option to play the audio file. Specifically, the user clicks the received link and plays it using the device's built-in player. This playback action allows the user to listen to the summarized content in audio.

[0156] Step 9:

[0157] User: Adds books they want to read to their favorites list. Specifically, they press the "Add to Favorites" button on their device. This action saves the book's information to the user's personal data store.

[0158] Step 10:

[0159] Server: When a new summary is generated for a book added to the favorites list, a notification is sent. Specifically, when new summary data is generated, the user is notified via email or app notification. This notification allows the user to check the new information immediately.

[0160] (Application Example 1)

[0161] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0162] For users leading busy lives, there is a need for efficient ways to understand the content of books. However, traditional methods of providing book summaries and audiobooks sometimes make it difficult for users to quickly obtain the information they are looking for. Furthermore, there is a lack of systems that easily recommend new reading material and provide easy access to desired reading. A new system is needed to address these challenges.

[0163] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0164] In this invention, the server includes means for the user to input information about a reading material, means for receiving the information about the reading material input by the user and searching for the reading material in a database, means for summarizing the entire text of the searched reading material, means for converting the summarized content of the reading material into an audio file, means for distributing the audio file to the user's terminal device, and means for presenting recommended reading material based on the information about the reading material input by the user. As a result, the user can not only quickly listen to a summary in audio format based on the information they have input, but also receive recommendations for related new reading material, enabling them to understand the content of books more easily and efficiently.

[0165] A "user" is an individual or organization that uses the system to input information about reading material and receives the retrieved summaries and audio files.

[0166] "Reading materials" is a general term for information media provided in text format, such as books, magazines, articles, and ebooks.

[0167] "Means of input" refers to methods of providing information to a system using keyboards, voice input, touchscreens, etc.

[0168] "Means of receiving" refers to the function by which a server retrieves information sent by a user.

[0169] "Searching methods" refer to the function of finding relevant reading material from a database based on the received information.

[0170] "Full text" refers to text data that corresponds to the entire text of the searched reading material.

[0171] "Methods of summarization" refer to the function of extracting important parts from the entire text and summarizing them concisely.

[0172] "Means of converting to audio files" refers to a function that converts summarized text into audio data.

[0173] "Terminal devices" refer to electronic devices such as smartphones, tablets, and computers.

[0174] "Means of distribution" refers to the function of sending the generated audio file to the user's terminal device.

[0175] "Means for suggesting recommended reading material" refers to a function that suggests new reading material related to the user's input information.

[0176] The "playback option" refers to the choices a user can make to listen to an audio file.

[0177] "Adding to a favorites list" refers to a function that allows users to save reading material to a list so they can easily access it later.

[0178] "Means of sending notifications" refers to a function that allows the server to notify users when a new summary is generated.

[0179] This invention is a system designed to enable users to efficiently understand the content of reading materials. This system consists of a server and terminal devices, allowing users to input book information and obtain summarized content in audio format. A specific embodiment of this system is described below.

[0180] Hardware and software to be used

[0181] Hardware:

[0182] Terminal devices: Smartphones, tablets, computers, smart glasses, head-mounted displays, etc.

[0183] Server: A high-performance computer with internet connectivity.

[0184] software:

[0185] Python programming language

[0186] requests library: Used for communication between servers and terminal devices.

[0187] gTTS (Google Text-to-Speech) library: Converts text to speech.

[0188] playsound library: Plays the generated audio file.

[0189] System operation

[0190] Book Search

[0191] The user inputs the title and ISBN code of the book they want to read through the interface of the terminal device, and sends that information to the server. For example, the user inputs the title of "a famous novel."

[0192] Book summary generation

[0193] The server searches the database based on the input information and retrieves data on matching reading materials. The retrieved complete text is then summarized using a natural language processing engine. This summarization process extracts the main points and important parts of the text and condenses them into a concise form.

[0194] Audio file generation and distribution

[0195] The gTTS library is used to convert the generated summary text into an audio file. The generated audio file is delivered from the server to the terminal device, and the user can understand the summary by listening to it. For example, the content of a summarized "famous novel" is delivered as an audio file.

[0196] Recommended reading material

[0197] Based on the information entered about the reading material, the server generates and presents relevant recommended reading materials to the user. This allows the user to discover new and interesting reading material.

[0198] Specific example

[0199] For example, if a user enters the title "I Am a Cat," the system searches the database for the book, summarizes the entire text, and converts it into an audio file. This audio file is then delivered to the user's terminal device, allowing them to quickly listen to the content of "I Am a Cat." In addition, "Other Works by Soseki" are presented as recommended reading material.

[0200] Example of a prompt

[0201] Please enter the book title: I Am a Cat

[0202] Thus, the system of the present invention is a system that can efficiently provide summaries of reading materials and improve the reading experience by combining natural language processing technology and speech synthesis technology.

[0203] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0204] Step 1:

[0205] The user inputs information about the book they are reading from a terminal device. The input information, such as the book title and ISBN code, is transmitted to the server through the terminal device's interface.

[0206] Input: Book title and ISBN code

[0207] Output: Search information sent to the server

[0208] Step 2:

[0209] The server searches the database based on the information it receives about the reading material. The database stores the complete text of numerous reading materials and associated metadata.

[0210] Input: Search information (book title and ISBN code)

[0211] Output: Data on matching reading materials

[0212] Step 3:

[0213] The server passes the entire text retrieved as search results to a natural language processing (NLP) engine, which then generates a summary. The NLP engine extracts the important parts and creates a compact summary text.

[0214] Input: Entire text

[0215] Output: Summary text

[0216] Step 4:

[0217] The server inputs the generated summary text into a speech synthesis engine (e.g., the gTTS library) and converts it into an audio file. The audio file format commonly used is MP3.

[0218] Input: Summary text

[0219] Output: Audio file (MP3 format)

[0220] Step 5:

[0221] The server saves the generated audio file to temporary storage and generates a download link. The server then sends this link to the terminal device.

[0222] Input: Audio file (MP3 format)

[0223] Output: Download link

[0224] Step 6:

[0225] The user receives a download link from the terminal device, retrieves the audio file, and plays it on the terminal device's player. The terminal device then displays notifications and playback options.

[0226] Input: Download link

[0227] Output: Playback of audio file

[0228] Step 7:

[0229] Users add books they want to read to their favorites list. The server notifies the user via email or app notification when a new summary is generated.

[0230] Input: Registration information for your favorites list

[0231] Output: Notification of new summary

[0232] Step 8:

[0233] The server generates recommended reading material based on the user's input and presents it to the terminal device. Similar genres and authors are considered in the recommendations.

[0234] Input: User input information

[0235] Output: Presentation of recommended reading material

[0236] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0237] ---

[0238] This invention provides a system that enables users to efficiently understand the content of books even in their busy lives. This system combines a function that automatically generates book summaries and provides them in audio format with an emotion engine that recognizes the user's emotions.

[0239] System Overview

[0240] This system consists of a server, terminals, and an emotion engine. Users can input book information using the terminals and receive summarized content in audio. Furthermore, the emotion engine can suggest appropriate books and adjust summaries based on the user's emotions.

[0241] 1. Book Selection

[0242] User: Enter the title or ISBN code of the book you want to read using the terminal's interface. For example, if the user wants to read "a famous novel," they would enter the title in the terminal's search bar.

[0243] 2. Searching for book data

[0244] Terminal: Sends the book information entered by the user to the server.

[0245] Server: Based on the received information, it searches the database and finds data for matching books. It then sends the found data back to the terminal. For example, it might send the full text and related metadata of a famous novel to the terminal.

[0246] 3. Automatic summary generation

[0247] Server: The full text of a book is passed to a natural language processing (NLP) engine, which extracts the important parts and generates a summary. For example, the NLP engine extracts the key plot points and characters from the full text of a famous novel and summarizes them concisely.

[0248] 4. Adjusting emotion recognition and summarization

[0249] Server: Uses an emotion engine that analyzes the user's facial expressions, voice tone, input text, etc., to recognize the user's current emotional state. For example, the emotion engine might detect a happy emotion from the user's voice tone.

[0250] Server: Based on recognized emotions, adjusts the content and tone of the summary and audio file. For example, the emotion engine changes the tone of the summary according to the user's emotions, adjusting it to emphasize positive content.

[0251] 5. Generating audio files

[0252] Server: Inputs the adjusted summary into the speech synthesis engine and converts it into audio data. Specifically, the speech synthesis engine converts the summary into an MP3 audio file with an emotion-appropriate tone.

[0253] 6. Audio distribution

[0254] Server: Saves the generated audio files to temporary storage and generates a download link.

[0255] Device: Presents the received download link to the user. For example, a notification might appear stating, "You can listen to an audio summary of a famous novel."

[0256] 7. Audio Playback

[0257] User: Select an audio playback option from the device's notifications or interface. For example, the user presses the "Play" button.

[0258] Device: Download the audio file from the download link and play the audio using the built-in player.

[0259] 8. Add to favorites and receive notifications

[0260] User: Adds books they want to read to their favorites list. For example, the user presses the "Add to Favorites" button.

[0261] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[0262] In this way, the system of the present invention enables users to efficiently understand the content of books even in their busy lives, and further provides a personalized experience through the emotion engine. Each processing step is automated, making it easy for users to access and use. While specific examples have been provided, the system is not limited to these and can be applied in a variety of ways.

[0263] The following describes the processing flow.

[0264] ---

[0265] Step 1:

[0266] User: Open the dedicated application or webpage on your device and enter the title or ISBN code of the book you want to read. For example, if the user wants to read "a famous novel," they would enter the title in the search bar on their device.

[0267] Step 2:

[0268] Terminal: Sends the book information entered by the user to the server.

[0269] Step 3:

[0270] Server: Based on the title or ISBN code of the received book, the server searches the database for data on matching books.

[0271] Step 4:

[0272] Server: Returns the full text of the searched book and related metadata to the terminal.

[0273] Step 5:

[0274] Server: The full text of a book is passed to a natural language processing (NLP) engine, which extracts the important parts and generates a summary. For example, the NLP engine extracts the key plot points and characters from the full text of a famous novel and summarizes them concisely.

[0275] Step 6:

[0276] Server: After the summary is generated, an emotion engine analyzes the user's facial expressions, voice tone, input text, etc., to recognize the user's emotional state. For example, the emotion engine might detect a cheerful emotion from the user's voice tone.

[0277] Step 7:

[0278] Server: Based on the recognized emotion, adjust the summary and the content and tone of the voice file. For example, if the emotion engine determines that the user is in a happy mood, the tone of the summary is changed to positive.

[0279] Step 8:

[0280] Server: Input the adjusted summary into the text-to-speech engine and convert it into voice data. Specifically, the text-to-speech engine converts the summary content into an MP3 voice file with a tone corresponding to the emotion.

[0281] Step 9:

[0282] Server: Save the generated voice file in temporary storage and generate a download link.

[0283] Step 10:

[0284] Server: Send the download link to the terminal.

[0285] Step 11:

[0286] Terminal: Receive the sent download link and display a notification to the user saying "The summary can be played as voice".

[0287] Step 12:

[0288] User: Select the voice playback option in the terminal notification or interface and press the "Play" button.

[0289] Step 13:

[0290] Terminal: Retrieve the voice file from the download link and play the voice using the built-in player.

[0291] Step 14:

[0292] User: Add books you want to read to your favorites list as needed. For example, the user presses the "Add to Favorites" button.

[0293] Step 15:

[0294] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[0295] ---

[0296] The above is a detailed explanation of the processing steps. This system automates the entire process from book information input to audio playback, and further provides a personalized experience through its emotion engine. Users can easily access and use it.

[0297] (Example 2)

[0298] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0299] In modern society, users are required to consume information and content efficiently amidst their busy lives. However, traditional methods present challenges such as the lack of time to read long texts like books and the effort required to understand summarized information. Furthermore, providing personalized information tailored to the user's current emotional state is difficult.

[0300] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting information of content that the user wants to read, means for receiving the information of content input by the user and searching for the content from a database, means for automatically summarizing the full text of the searched content, means for adjusting the summarized content based on the user's emotional state, means for converting the adjusted content into an audio file, and means for delivering the audio file to the user's device. As a result, even in a busy life, the user can efficiently understand the main points of the content and receive personalized information tailored to their emotional state.

[0301] "User" refers to an individual or legal entity that intends to use this system to search for, summarize, or convert content into audio.

[0302] "Content" refers to information such as books, articles, and reports, and includes informational materials that users wish to read.

[0303] "Means of inputting information" refers to the interface that allows users to input information such as the content title, author name, and ISBN code into their device.

[0304] A "database" refers to a collection of information that stores, manages, and allows for quick searching of content metadata and full text.

[0305] "Methods for automatic summarization" refer to algorithms or engines that use natural language processing technology to extract key parts from the full text of content and generate a concise summary.

[0306] "Means of adjustment based on emotional state" refers to a function that analyzes the user's current emotions and adjusts the content and tone of information such as summaries and audio according to those emotions.

[0307] The "means for converting to an audio file" refers to a text-to-speech engine or software for converting automatically summarized text information into audio data.

[0308] The "device" refers to an electronic device for a user to use this system, including smartphones, tablets, personal computers, etc.

[0309] The "means for distributing" refers to the function of transmitting the generated audio file through a network to provide it to the user.

[0310] The present invention is a system provided for a user to efficiently understand the content of the content even in a busy life. This system combines the function of automatically generating a summary of the content and providing it in audio with an emotion engine that recognizes the user's emotions.

[0311] System Configuration

[0312] This system is composed of a server, a terminal, and an emotion engine. A user can input content information using the terminal and obtain the summarized content in audio. Furthermore, the emotion engine can propose appropriate content and adjust the summary according to the user's emotions.

[0313] Hardware and Software to be Used

[0314] 1. Server: Use a cloud server or a physical server with high computing power.

[0315] Database: Use a database solution such as AWS (registered trademark) RDS or MongoDB.

[0316] Natural Language Processing Engine: Use a generative AI model such as OpenAI (registered trademark) GPT-3.

[0317] Emotion recognition engine: Uses emotion recognition services such as Microsoft® Azure® Emotion API.

[0318] Text-to-speech engine: Uses a text-to-speech service such as the Google Text-to-Speech API.

[0319] 2. Device: A smartphone, tablet, or personal computer used by the user.

[0320] User interface: Web browser or dedicated application.

[0321] Program processing flow

[0322] The server receives content information entered by the user, searches the database based on that information, and retrieves the full text of the relevant content. The retrieved full text is then passed to a natural language processing engine, which automatically generates a summary. Furthermore, based on the user's sentiment data obtained from the device, an emotion recognition engine is used to adjust the summary content and voice tone. The adjusted summary is input to a speech synthesis engine and generated as an MP3 audio file. This audio file is stored in temporary storage by the server, and a download link is generated. Finally, the link is sent to the user's device, and the user can play the audio file from this link.

[0323] Specific example

[0324] For example, if a user wants to read a famous novel, they enter the title into the search bar on their device. This information is sent to the server, which retrieves the full text of the novel from its database. The retrieved text is summarized by a generative AI model. Based on this summary, an emotion recognition engine recognizes the user's emotions from their facial expressions and voice tone, and adjusts the content and tone of the summary. The adjusted summary is converted into an MP3 audio file by a speech synthesis engine and provided to the user's device.

[0325] Example of a prompt

[0326] "Please provide an audio summary of Natsume Soseki's 'Kokoro,' with tone adjustments based on emotional recognition."

[0327] "Please automatically summarize the latest business books and generate audio files with tone adjustments based on emotion recognition."

[0328] In this way, the system of the present invention enables users to efficiently understand the key points of content even in their busy lives and to provide personalized information tailored to their emotional state.

[0329] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0330] Step 1: Enter content information

[0331] User: Enter the title or ISBN code of the content you want to read from the terminal's interface. For example, if the user wants to read "a famous novel," they would enter the title in the terminal's search bar.

[0332] Input: Title and ISBN code entered by the user.

[0333] Output: Input data is sent to the server.

[0334] Specific operation: The terminal interface displays "Please enter the title or ISBN of the content." The user enters the information using the keyboard and then clicks the "Search" button.

[0335] Step 2: Search for content data

[0336] Terminal: Sends information about the content entered by the user to the server.

[0337] Server: Based on the received information, it searches a database (e.g., AWS RDS or MongoDB) and finds data with matching content. It then sends the corresponding data back to the terminal in JSON format.

[0338] Input: Entered content information.

[0339] Output: The full text and associated metadata of the relevant content retrieved from the database.

[0340] Specific operation: The terminal sends input data to the server using a REST API. The server executes SQL or NoSQL queries against the database, retrieves the relevant data, converts it to JSON format, and sends it back to the terminal.

[0341] Step 3: Generate an automated summary

[0342] Server: The full text of the content is passed to a natural language processing (NLP) engine (e.g., OpenAI GPT-3 model), which extracts the important parts and generates a summary.

[0343] Input: Full text (JSON format).

[0344] Output: Summarized text.

[0345] Specific operation: The server inputs the full text of the retrieved content into the AI ​​model along with the prompt, "Please summarize:". The AI ​​model analyzes the main parts of the text, extracts the important parts, and generates a summary. The generated summary is saved as a text file.

[0346] Step 4: Adjusting Emotion Recognition and Summarization

[0347] User: While generating a summary, the device's built-in camera and microphone are used to collect user emotion data. For example, when the user is smiling while looking at the screen.

[0348] Server: Uses an emotion engine (e.g., Microsoft Azure Emotion API) that analyzes the user's facial expressions, voice tone, input text, etc., to recognize the user's current emotional state.

[0349] Input: User's facial expression, voice tone, and input text.

[0350] Output: Recognized emotion data.

[0351] Specific operation: The user's device captures facial expressions with its camera and collects audio with its microphone. This data is sent to a server, which uses an emotion engine to send a prompt message saying, "Please recognize the emotion." Based on the recognized emotion data, the content and tone of the summary are adjusted.

[0352] Step 5: Generate audio files

[0353] Server: Inputs the refined summary into a text-to-speech engine (e.g., Google Text-to-Speech API) and converts it into audio data.

[0354] Input: Adjusted summary text.

[0355] Output: MP3 audio file.

[0356] Specific operation: The server inputs the adjusted summary into the speech synthesis engine, along with the prompt, "Please convert this text into speech, including emotional tone." The generated audio data is saved to the server as an MP3 file.

[0357] Step 6: Audio Distribution

[0358] Server: Saves the generated audio files to temporary storage (e.g., Amazon S3) and generates a download link.

[0359] Terminal: Presents the received download link to the user.

[0360] Input: MP3 audio file.

[0361] Output: Download link.

[0362] Specific operation: The server uploads the audio file to temporary storage and generates a public download link. This link is sent to the device, and the device notifies the user of the link.

[0363] Step 7: Play back audio

[0364] User: Select an audio playback option from the device's notifications or interface. For example, the user presses the "Play" button.

[0365] Device: Download the audio file from the download link and play the audio using the built-in player.

[0366] Input: Download link.

[0367] Output: Audio played on the device.

[0368] Specific operation: When the user presses the "Play" button, the device retrieves the audio file from the download link, launches the media player, and plays the audio.

[0369] Step 8: Add to favorites and receive notifications

[0370] User: Adds content they want to read to their favorites list. For example, the user presses the "Add to Favorites" button.

[0371] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[0372] Input: Information about the content to add to your favorites list.

[0373] Output: Notification when a new summary is generated.

[0374] Specific operation: When a user presses the "Add to Favorites" button, information about that content is sent to the server and stored in the database. The server sends an email or app notification to the user when a new summary is added.

[0375] (Application Example 2)

[0376] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0377] In modern society, users lead busy lives and find it difficult to dedicate sufficient time to reading books. Therefore, there is a need for a system that allows for efficient comprehension of book content. However, existing systems merely provide summaries and lack personalization based on user emotions, resulting in low satisfaction. Furthermore, the audio summaries do not adapt to the user's state of mind, leading to insufficient information comprehension and engagement.

[0378] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting information of a book that the user wants to read, means for receiving the book information input by the user and searching for the book in a database, means for summarizing the full text of the searched book, means for converting the summarized book content into an audio file, means for distributing the audio file to the user's terminal, means for recognizing the user's emotions, and means for adjusting the summary content and tone based on the recognized emotions. As a result, users can efficiently understand the content of books even in their busy lives, and receive personalized book summaries based on their emotions in audio format, enabling a more satisfying information provision.

[0379] "User" refers to a person who uses the system to obtain summary information about books.

[0380] "Book information" refers to data used to identify a book, such as the book's title, ISBN code, and author's name.

[0381] A "database" refers to a collection of information that stores the full text of a book and related metadata.

[0382] "Full text" refers to the entire body of text in a book.

[0383] A "summary" refers to information that has been simplified by extracting the most important parts from the full text of a book.

[0384] An "audio file" refers to a digital file in which the content of a summarized book has been converted into audio data.

[0385] "Device" refers to electronic devices used by users, such as smartphones, smart glasses, and head-mounted displays.

[0386] "Emotion recognition" refers to technology that detects and analyzes the emotional state of a user.

[0387] "Tone" refers to the quality or tone of voice, and means a characteristic of voice that is altered based on emotion.

[0388] "Distribution" refers to the act of sending data from a server to a terminal.

[0389] "Personalization" refers to customization tailored to the user's specific needs and preferences.

[0390] Modes for carrying out the invention

[0391] System Overview

[0392] The system of this invention consists of a server, a terminal, and an emotion engine that recognizes the user's emotions. The user can use the terminal to input information about a specific book and obtain a summary of that book as an audio file. Furthermore, the emotion engine is used to adjust the summary according to the user's emotions.

[0393] Hardware and software used

[0394] Server: AWS (Amazon Web Services) or GCP (Google Cloud Platform)

[0395] Database: PostgreSQL or MySQL

[0396] NLP engine: SpaCy or BERT, both Python libraries.

[0397] Emotion Engine: An emotion recognition model using OpenCV and TENSORFLOW®.

[0398] Text-to-speech engine: Amazon Polly or Google Text-to-Speech

[0399] Devices: Smartphones, smart glasses, head-mounted displays

[0400] Explanation of the program's processing

[0401] 1. Book Selection

[0402] The user enters the title or ISBN code of the book they want to read through the terminal's interface. For example, if the user wants to read the book "1984," they would enter "1984" into the terminal's search bar.

[0403] 2. Searching for book data

[0404] The server receives the book information entered by the user, searches its database, and identifies the full text of any matching books. It then sends this result back to the terminal.

[0405] 3. Automatic summary generation

[0406] The server uses an NLP engine (e.g., spaCy or BERT) to extract key parts from the full text of the searched books and generate summaries.

[0407] 4. Emotion recognition

[0408] The device uses its camera and microphone to capture the user's facial expressions and voice tone, which are then analyzed by an emotion engine (e.g., OpenCV and TensorFlow models).

[0409] 5. Summarizing and adjusting the tone

[0410] Based on the analyzed sentiment data, the server adjusts the summary content and voice tone to personalize each user. This tone adjustment is performed during the voice file generation process.

[0411] 6. Generating audio files

[0412] The server uses a text-to-speech engine (e.g., Amazon Polly or Google Text-to-Speech) to convert the refined summary into audio data.

[0413] 7. Audio distribution

[0414] The server uploads the generated audio files to cloud storage and provides a download link to the user's device. The user can then retrieve the audio files via this link and make them available for use.

[0415] Specific example

[0416] The user opens the application and types "1984". The camera captures the user's face and determines that they are relaxed. The server uses this emotional information to generate a summary of the book "1984" in a friendly tone and provides it as an audio file. This entire process allows the user to efficiently understand the book's content even in a busy environment.

[0417] Example of a prompt

[0418] "Summarize the key information from all the texts from 1984 in 50 characters or less."

[0419] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0420] Step 1:

[0421] The user enters the title and ISBN code of the book they want to read through the terminal's interface. For example, if the user wants to read "1984," they enter "1984" into the terminal's search bar. This action results in data transfer, where the terminal sends the entered book information to the server.

[0422] Step 2:

[0423] The server queries the database for book information received from the terminal and searches for matching book data (full text and metadata). During this process, the server uses the title and ISBN code as keys to execute the search query. It retrieves the full text of the relevant books as a search result and sends it back to the terminal.

[0424] Step 3:

[0425] The server passes the full text of the book, retrieved from the database, to a natural language processing (NLP) engine. The NLP engine analyzes the text, extracts key parts, and generates a summary. For example, the NLP model might create a summary of the book's main plot and characters in 50 characters or less. This summary is temporarily stored.

[0426] Step 4:

[0427] The device uses its built-in camera and microphone to capture the user's facial expressions and voice tone. The collected data is sent to a server in real time. For example, the camera can capture the user's smile and send it as input data to an emotion recognition model.

[0428] Step 5:

[0429] The server uses an emotion recognition model (e.g., OpenCV and TensorFlow) to analyze the user's emotional state. It recognizes the user's emotions (e.g., joy, sadness, surprise, etc.) from the video and audio provided as input data. This analysis result is used to refine the summary data as an emotional state.

[0430] Step 6:

[0431] The server adjusts the content and tone of the summary generated by the NLP engine based on the user's emotional state. For example, if the user's emotion is "relaxed," the summary's tone will be adjusted to softer language to make it more approachable. As a result of this step, the adjusted summary data is generated.

[0432] Step 7:

[0433] The server inputs the adjusted summary data into a text-to-speech engine (e.g., Amazon Polly or Google Text-to-Speech) to generate an audio file. The text-to-speech engine generates audio data that matches the summary content and emotional tone, and converts it into an MP3 audio file. This audio file is stored on the server.

[0434] Step 8:

[0435] The server uploads the generated audio file to cloud storage and generates a download link. This link is resent to the device, and the user receives a notification. By clicking the link, the user can download and listen to the audio file.

[0436] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0437] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0438] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0439] [Second Embodiment]

[0440] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0441] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0442] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0443] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0444] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0445] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0446] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0447] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0448] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0449] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0450] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0451] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0452] ---

[0453] This invention provides a system that enables users to efficiently understand the content of books even in their busy lives. This system has the function of automatically generating a book summary and providing it in audio format. The program processing of this system is described in detail below.

[0454] System Overview

[0455] This system consists of a server and terminals, allowing users to input book information via the terminals and receive summarized content in audio format. The system has the following main functions:

[0456] 1. Book Selection

[0457] User: Enter the title or ISBN code of the book you want to read using the terminal's interface. For example, if the user wants to read "a famous novel," they would enter the title in the terminal's search bar.

[0458] 2. Searching for book data

[0459] Terminal: Sends the book information entered by the user to the server.

[0460] Server: Based on the received information, it searches the database and finds data for matching books. It then sends the found data back to the terminal. For example, it might send the full text and related metadata of a famous novel to the terminal.

[0461] 3. Automatic summary generation

[0462] Server: The full text of a book is passed to a natural language processing (NLP) engine, which extracts the important parts and generates a short summary. For example, the NLP engine extracts the key plot points and characters from the full text of a famous novel and summarizes them concisely.

[0463] 4. Generating audio files

[0464] Server: Inputs the generated summary into the speech synthesis engine and converts it into audio data. For example, the speech synthesis engine converts the summary content into an MP3 audio file.

[0465] 5. Audio distribution

[0466] Server: Saves the audio file to temporary storage and generates a download link.

[0467] Device: Presents the received download link to the user. For example, a notification might appear stating, "You can listen to an audio summary of a famous novel."

[0468] 6. Audio Playback

[0469] User: Select an audio playback option from the device's notifications or interface. For example, the user presses the "Play" button.

[0470] Device: Download the audio file from the download link and play it using the built-in player.

[0471] 7. Add to favorites and receive notifications

[0472] User: Adds books they want to read to their favorites list. For example, the user presses the "Add to Favorites" button.

[0473] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[0474] In this way, the system of the present invention provides users with a means to efficiently understand the content of books even in their busy lives. By automating each processing step, users can easily access and use the system. Although specific examples have been provided, the system is not limited to these and can be applied in a variety of ways.

[0475] The following describes the processing flow.

[0476] ---

[0477] Step 1:

[0478] User: Open the dedicated application or webpage on your device and enter the title or ISBN code of the book you want to read.

[0479] Step 2:

[0480] Terminal: Sends the book information entered by the user to the server.

[0481] Step 3:

[0482] Server: Based on the title or ISBN code of the received book, the server searches the database for data on matching books.

[0483] Step 4:

[0484] Server: Returns the full text of the searched book and related metadata to the terminal.

[0485] Step 5:

[0486] Server: Passes the full text to a natural language processing (NLP) engine, which extracts the important parts and generates a summary.

[0487] Step 6:

[0488] Server: Inputs the generated summary into the speech synthesis engine and converts it into audio data. Specifically, the speech synthesis engine converts the summary content into an MP3 audio file.

[0489] Step 7:

[0490] Server: Saves the generated audio files to temporary storage and generates a download link.

[0491] Step 8:

[0492] Server: Sends a download link to the device.

[0493] Step 9:

[0494] Terminal: Receives the sent download link and displays a notification to the user stating, "The summary is available for audio playback."

[0495] Step 10:

[0496] User: Select the audio playback option in the device's notification or interface and press the "Play" button.

[0497] Step 11:

[0498] Device: Download the audio file from the download link and play the audio using the built-in player.

[0499] Step 12:

[0500] User: Add books you want to read to your favorites list as needed. For example, the user presses the "Add to Favorites" button.

[0501] Step 13:

[0502] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[0503] ---

[0504] The above is a detailed explanation of the processing steps. This system automates the entire process, from inputting book information to audio playback, to help users efficiently understand the content of books.

[0505] (Example 1)

[0506] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0507] In today's busy lives, many people don't have the time to efficiently understand the content of books. There is a growing demand for systems that allow users to easily access summarized book content and listen to it as audio. Traditional systems lacked convenience because the processes of generating summaries, converting to audio files, and distribution were time-consuming and laborious.

[0508] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0509] In this invention, the server includes means for inputting information about a book that the user wants to read, means for receiving the book information input by the user and searching for the book in a data store, means for summarizing the entire text of the searched book, means for inputting the generated summary into a natural language processing engine, means for inputting the generated summary into a speech synthesis engine and converting it into audio data, means for temporarily storing the audio data and generating a download link thereof, and means for delivering the download link to the user's terminal. This makes it possible for the user to efficiently and quickly obtain an audio summary of the book.

[0510] A "user" is an individual or organization that uses the system to obtain book summaries and audio data.

[0511] "Book information" refers to information necessary to uniquely identify a book, such as the book's title, author's name, and ISBN code.

[0512] A "data store" is a storage device or database system used to store and manage book data.

[0513] "The full text of the book" refers to all the text and data contained in the book.

[0514] A "summary" refers to a concise compilation of the most important parts and key points extracted from the entire text of a book.

[0515] A "natural language processing engine" is software or algorithms that analyze text data and perform tasks such as summarizing, translating, and other natural language processing.

[0516] A "speech synthesis engine" is a software or hardware system that takes text data as input and converts it into speech data.

[0517] "Audio data" refers to audio files in an audible format, generated by a speech synthesis engine.

[0518] "Temporary storage" refers to the process of saving data only for a certain period, with deletion or updating performed as needed.

[0519] A "download link" refers to a URL or path used to obtain audio data or other data.

[0520] A "terminal" is an electronic device used by a user to access and operate a system.

[0521] This invention provides a system that enables users to efficiently understand the content of books even in their busy lives. This system has the function of automatically generating a book summary and providing it in audio format. Specific embodiments of this system are described in detail below.

[0522] System Configuration

[0523] This system primarily consists of a server and terminals, and users obtain book summaries as audio via the terminals. The system is comprised of the following hardware and software combinations.

[0524] hardware

[0525] The server includes a data store (e.g., a MySQL database), an NLP engine, and a speech synthesis engine.

[0526] Device: An electronic device used by a user, such as a smartphone, tablet, or personal computer.

[0527] software

[0528] Natural language processing engine (e.g., GPT-3)

[0529] Text-to-speech engine (e.g., Google Text-to-Speech)

[0530] Database management system (e.g., MySQL)

[0531] HTTP request processing libraries (e.g., Axios)

[0532] Operation flow

[0533] The operation flow of this system is as follows:

[0534] 1. Book Selection

[0535] The user enters the title or ISBN code of the book they want to read through the terminal's interface. For example, if the user wants to read "a famous novel," they enter the title in the terminal's search bar and press the "Search" button.

[0536] 2. Searching for book data

[0537] The terminal sends the book information entered by the user to the server. Based on the received information, the server searches its data store for matching book data. As a result, it sends back the full text of "a famous novel" and related metadata to the terminal.

[0538] 3. Automatic summary generation

[0539] The server passes the full text of the book to a natural language processing engine, which extracts the important parts and generates a short summary. For example, it passes a prompt sentence like the following to the natural language processing engine:

[0540] Please summarize the following text: "The complete text of a famous novel."

[0541] 4. Generating audio files

[0542] The server inputs the generated summary into the speech synthesis engine and converts it into speech data. Specifically, it uses prompts like the following:

[0543] "Please convert this summary into an audio file."

[0544] The generated audio file will be temporarily saved.

[0545] 5. Audio distribution

[0546] The server generates a download link for the audio file and notifies the terminal of that link. For example, it sends the user a message saying, "You can listen to an audio summary of a famous novel." The terminal then displays the received download link on its interface.

[0547] 6. Audio Playback

[0548] The user selects the audio playback option through the device's notifications or interface. When the user presses the "Play" button, the device retrieves the audio file from the download link and plays it using the built-in player.

[0549] 7. Add to favorites and receive notifications

[0550] Users can add books they want to read to their favorites list. When a new summary is generated, the server notifies the user via email or app notification.

[0551] This system enables users to efficiently understand the content of books even in their busy lives, and by automating each processing step, it achieves simple access and use. For example, it can be used in the same process when a user is trying to read the latest business book, "The Road to Success."

[0552] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0553] Step 1:

[0554] User: Enter information about the book you want to read. Specifically, the user enters the book title and ISBN code through the terminal interface and presses the "Search" button. This entered information is then sent to the next processing step.

[0555] Step 2:

[0556] Terminal: Sends the book information entered by the user to the server. Specifically, it sends the book information to the server using an HTTP request. This operation passes the input data to the server.

[0557] Step 3:

[0558] Server: Searches the data store based on the received book information. Specifically, it queries the data store (e.g., a MySQL database) based on the received book information and retrieves the full text and metadata of matching books. This query process outputs the full text and metadata.

[0559] Step 4:

[0560] Server: The server passes the full text obtained as search results to a natural language processing (NLP) engine to generate a summary. Specifically, it inputs the full text into the NLP engine using the prompt message, "Summarize the following text: The full text of 'Book Title'." The NLP engine performs text analysis, extracts the important parts, and generates a summary. This summary is then passed to the next step.

[0561] Step 5:

[0562] Server: Inputs the generated summary into a text-to-speech engine and converts it into audio data. Specifically, it inputs the summary into a text-to-speech engine (e.g., Google Text-to-Speech) using a prompt message such as "Please convert this summary into an audio file." The text-to-speech engine converts the text into speech and generates an audio file in MP3 format. This audio file is then passed on to the next step.

[0563] Step 6:

[0564] Server: Temporarily stores the audio file and generates a download link. Specifically, it uploads the generated audio file to cloud storage and generates a download link. This link is then sent to the next step.

[0565] Step 7:

[0566] Device: Notifies the user of the generated download link. Specifically, the device's interface and notification system display the link along with the message, "You can listen to an audio summary of the book title." When the user clicks this link, they proceed to the next step.

[0567] Step 8:

[0568] User: Selects the option to play the audio file. Specifically, the user clicks the received link and plays it using the device's built-in player. This playback action allows the user to listen to the summarized content in audio.

[0569] Step 9:

[0570] User: Adds books they want to read to their favorites list. Specifically, they press the "Add to Favorites" button on their device. This action saves the book's information to the user's personal data store.

[0571] Step 10:

[0572] Server: When a new summary is generated for a book added to the favorites list, a notification is sent. Specifically, when new summary data is generated, the user is notified via email or app notification. This notification allows the user to check the new information immediately.

[0573] (Application Example 1)

[0574] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0575] For users leading busy lives, there is a need for efficient ways to understand the content of books. However, traditional methods of providing book summaries and audiobooks sometimes make it difficult for users to quickly obtain the information they are looking for. Furthermore, there is a lack of systems that easily recommend new reading material and provide easy access to desired reading. A new system is needed to address these challenges.

[0576] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0577] In this invention, the server includes means for the user to input information about a reading material, means for receiving the information about the reading material input by the user and searching for the reading material in a database, means for summarizing the entire text of the searched reading material, means for converting the summarized content of the reading material into an audio file, means for distributing the audio file to the user's terminal device, and means for presenting recommended reading material based on the information about the reading material input by the user. As a result, the user can not only quickly listen to a summary in audio format based on the information they have input, but also receive recommendations for related new reading material, enabling them to understand the content of books more easily and efficiently.

[0578] A "user" is an individual or organization that uses the system to input information about reading material and receives the retrieved summaries and audio files.

[0579] "Reading materials" is a general term for information media provided in text format, such as books, magazines, articles, and ebooks.

[0580] "Means of input" refers to methods of providing information to a system using keyboards, voice input, touchscreens, etc.

[0581] "Means of receiving" refers to the function by which a server retrieves information sent by a user.

[0582] "Searching methods" refer to the function of finding relevant reading material from a database based on the received information.

[0583] "Full text" refers to text data that corresponds to the entire text of the searched reading material.

[0584] "Methods of summarization" refer to the function of extracting important parts from the entire text and summarizing them concisely.

[0585] "Means of converting to audio files" refers to a function that converts summarized text into audio data.

[0586] "Terminal devices" refer to electronic devices such as smartphones, tablets, and computers.

[0587] "Means of distribution" refers to the function of sending the generated audio file to the user's terminal device.

[0588] "Means for suggesting recommended reading material" refers to a function that suggests new reading material related to the user's input information.

[0589] The "playback option" refers to the choices a user can make to listen to an audio file.

[0590] "Adding to a favorites list" refers to a function that allows users to save reading material to a list so they can easily access it later.

[0591] "Means of sending notifications" refers to a function that allows the server to notify users when a new summary is generated.

[0592] This invention is a system designed to enable users to efficiently understand the content of reading materials. This system consists of a server and terminal devices, allowing users to input book information and obtain summarized content in audio format. A specific embodiment of this system is described below.

[0593] Hardware and software to be used

[0594] Hardware:

[0595] Terminal devices: Smartphones, tablets, computers, smart glasses, head-mounted displays, etc.

[0596] Server: A high-performance computer with internet connectivity.

[0597] software:

[0598] Python programming language

[0599] requests library: Used for communication between servers and terminal devices.

[0600] gTTS (Google Text-to-Speech) library: Converts text to speech.

[0601] playsound library: Plays the generated audio file.

[0602] System operation

[0603] Book Search

[0604] The user inputs the title and ISBN code of the book they want to read through the interface of the terminal device, and sends that information to the server. For example, the user inputs the title of "a famous novel."

[0605] Book summary generation

[0606] The server searches the database based on the input information and retrieves data on matching reading materials. The retrieved complete text is then summarized using a natural language processing engine. This summarization process extracts the main points and important parts of the text and condenses them into a concise form.

[0607] Audio file generation and distribution

[0608] The gTTS library is used to convert the generated summary text into an audio file. The generated audio file is delivered from the server to the terminal device, and the user can understand the summary by listening to it. For example, the content of a summarized "famous novel" is delivered as an audio file.

[0609] Recommended reading material

[0610] Based on the information entered about the reading material, the server generates and presents relevant recommended reading materials to the user. This allows the user to discover new and interesting reading material.

[0611] Specific example

[0612] For example, if a user enters the title "I Am a Cat," the system searches the database for the book, summarizes the entire text, and converts it into an audio file. This audio file is then delivered to the user's terminal device, allowing them to quickly listen to the content of "I Am a Cat." In addition, "Other Works by Soseki" are presented as recommended reading material.

[0613] Example of a prompt

[0614] Please enter the book title: I Am a Cat

[0615] Thus, the system of the present invention is a system that can efficiently provide summaries of reading materials and improve the reading experience by combining natural language processing technology and speech synthesis technology.

[0616] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0617] Step 1:

[0618] The user inputs information about the book they are reading from a terminal device. The input information, such as the book title and ISBN code, is transmitted to the server through the terminal device's interface.

[0619] Input: Book title and ISBN code

[0620] Output: Search information sent to the server

[0621] Step 2:

[0622] The server searches the database based on the information it receives about the reading material. The database stores the complete text of numerous reading materials and associated metadata.

[0623] Input: Search information (book title and ISBN code)

[0624] Output: Data on matching reading materials

[0625] Step 3:

[0626] The server passes the entire text retrieved as search results to a natural language processing (NLP) engine, which then generates a summary. The NLP engine extracts the important parts and creates a compact summary text.

[0627] Input: Entire text

[0628] Output: Summary text

[0629] Step 4:

[0630] The server inputs the generated summary text into a speech synthesis engine (e.g., the gTTS library) and converts it into an audio file. The audio file format commonly used is MP3.

[0631] Input: Summary text

[0632] Output: Audio file (MP3 format)

[0633] Step 5:

[0634] The server saves the generated audio file to temporary storage and generates a download link. The server then sends this link to the terminal device.

[0635] Input: Audio file (MP3 format)

[0636] Output: Download link

[0637] Step 6:

[0638] The user receives a download link from the terminal device, retrieves the audio file, and plays it on the terminal device's player. The terminal device then displays notifications and playback options.

[0639] Input: Download link

[0640] Output: Playback of audio file

[0641] Step 7:

[0642] Users add books they want to read to their favorites list. The server notifies the user via email or app notification when a new summary is generated.

[0643] Input: Registration information for your favorites list

[0644] Output: Notification of new summary

[0645] Step 8:

[0646] The server generates recommended reading material based on the user's input and presents it to the terminal device. Similar genres and authors are considered in the recommendations.

[0647] Input: User input information

[0648] Output: Presentation of recommended reading material

[0649] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0650] ---

[0651] This invention provides a system that enables users to efficiently understand the content of books even in their busy lives. This system combines a function that automatically generates book summaries and provides them in audio format with an emotion engine that recognizes the user's emotions.

[0652] System Overview

[0653] This system consists of a server, terminals, and an emotion engine. Users can input book information using the terminals and receive summarized content in audio. Furthermore, the emotion engine can suggest appropriate books and adjust summaries based on the user's emotions.

[0654] 1. Book Selection

[0655] User: Enter the title or ISBN code of the book you want to read using the terminal's interface. For example, if the user wants to read "a famous novel," they would enter the title in the terminal's search bar.

[0656] 2. Searching for book data

[0657] Terminal: Sends the book information entered by the user to the server.

[0658] Server: Based on the received information, it searches the database and finds data for matching books. It then sends the found data back to the terminal. For example, it might send the full text and related metadata of a famous novel to the terminal.

[0659] 3. Automatic summary generation

[0660] Server: The full text of a book is passed to a natural language processing (NLP) engine, which extracts the important parts and generates a summary. For example, the NLP engine extracts the key plot points and characters from the full text of a famous novel and summarizes them concisely.

[0661] 4. Adjusting emotion recognition and summarization

[0662] Server: Uses an emotion engine that analyzes the user's facial expressions, voice tone, input text, etc., to recognize the user's current emotional state. For example, the emotion engine might detect a happy emotion from the user's voice tone.

[0663] Server: Based on recognized emotions, adjusts the content and tone of the summary and audio file. For example, the emotion engine changes the tone of the summary according to the user's emotions, adjusting it to emphasize positive content.

[0664] 5. Generating audio files

[0665] Server: Inputs the adjusted summary into the speech synthesis engine and converts it into audio data. Specifically, the speech synthesis engine converts the summary into an MP3 audio file with an emotion-appropriate tone.

[0666] 6. Audio distribution

[0667] Server: Saves the generated audio files to temporary storage and generates a download link.

[0668] Device: Presents the received download link to the user. For example, a notification might appear stating, "You can listen to an audio summary of a famous novel."

[0669] 7. Audio Playback

[0670] User: Select an audio playback option from the device's notifications or interface. For example, the user presses the "Play" button.

[0671] Device: Download the audio file from the download link and play the audio using the built-in player.

[0672] 8. Add to favorites and receive notifications

[0673] User: Adds books they want to read to their favorites list. For example, the user presses the "Add to Favorites" button.

[0674] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[0675] In this way, the system of the present invention enables users to efficiently understand the content of books even in their busy lives, and further provides a personalized experience through the emotion engine. Each processing step is automated, making it easy for users to access and use. While specific examples have been provided, the system is not limited to these and can be applied in a variety of ways.

[0676] The following describes the processing flow.

[0677] ---

[0678] Step 1:

[0679] User: Open the dedicated application or webpage on your device and enter the title or ISBN code of the book you want to read. For example, if the user wants to read "a famous novel," they would enter the title in the search bar on their device.

[0680] Step 2:

[0681] Terminal: Sends the book information entered by the user to the server.

[0682] Step 3:

[0683] Server: Based on the title or ISBN code of the received book, the server searches the database for data on matching books.

[0684] Step 4:

[0685] Server: Returns the full text of the searched book and related metadata to the terminal.

[0686] Step 5:

[0687] Server: The full text of a book is passed to a natural language processing (NLP) engine, which extracts the important parts and generates a summary. For example, the NLP engine extracts the key plot points and characters from the full text of a famous novel and summarizes them concisely.

[0688] Step 6:

[0689] Server: After the summary is generated, an emotion engine analyzes the user's facial expressions, voice tone, input text, etc., to recognize the user's emotional state. For example, the emotion engine might detect a cheerful emotion from the user's voice tone.

[0690] Step 7:

[0691] Server: Adjusts the content and tone of the summary and audio file based on recognized emotions. For example, the emotion engine might change the tone of the summary to a positive one based on the user's positive emotions.

[0692] Step 8:

[0693] Server: Inputs the adjusted summary into the speech synthesis engine and converts it into audio data. Specifically, the speech synthesis engine converts the summary into an MP3 audio file with an emotion-appropriate tone.

[0694] Step 9:

[0695] Server: Saves the generated audio files to temporary storage and generates a download link.

[0696] Step 10:

[0697] Server: Sends a download link to the device.

[0698] Step 11:

[0699] Terminal: Receives the sent download link and displays a notification to the user stating, "The summary is available for audio playback."

[0700] Step 12:

[0701] User: Select the audio playback option in the device's notification or interface and press the "Play" button.

[0702] Step 13:

[0703] Device: Download the audio file from the download link and play the audio using the built-in player.

[0704] Step 14:

[0705] User: Add books you want to read to your favorites list as needed. For example, the user presses the "Add to Favorites" button.

[0706] Step 15:

[0707] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[0708] ---

[0709] The above is a detailed explanation of the processing steps. This system automates the entire process from book information input to audio playback, and further provides a personalized experience through its emotion engine. Users can easily access and use it.

[0710] (Example 2)

[0711] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0712] In modern society, users are required to consume information and content efficiently amidst their busy lives. However, traditional methods present challenges such as the lack of time to read long texts like books and the effort required to understand summarized information. Furthermore, providing personalized information tailored to the user's current emotional state is difficult.

[0713] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting information of content that the user wants to read, means for receiving the information of content input by the user and searching for the content from a database, means for automatically summarizing the full text of the searched content, means for adjusting the summarized content based on the user's emotional state, means for converting the adjusted content into an audio file, and means for delivering the audio file to the user's device. As a result, even in a busy life, the user can efficiently understand the main points of the content and receive personalized information tailored to their emotional state.

[0714] "User" refers to an individual or legal entity that intends to use this system to search for, summarize, or convert content into audio.

[0715] "Content" refers to information such as books, articles, and reports, and includes informational materials that users wish to read.

[0716] "Means of inputting information" refers to the interface that allows users to input information such as the content title, author name, and ISBN code into their device.

[0717] A "database" refers to a collection of information that stores, manages, and allows for quick searching of content metadata and full text.

[0718] "Methods for automatic summarization" refer to algorithms or engines that use natural language processing technology to extract key parts from the full text of content and generate a concise summary.

[0719] "Means of adjustment based on emotional state" refers to a function that analyzes the user's current emotions and adjusts the content and tone of information such as summaries and audio according to those emotions.

[0720] "Means of converting to audio files" refers to speech synthesis engines and software that convert automatically summarized text information into audio data.

[0721] "Device" refers to the electronic device that a user uses to access this system, and includes smartphones, tablets, personal computers, etc.

[0722] "Means of distribution" refers to the function of transmitting generated audio files over a network in order to provide them to users.

[0723] This invention is a system designed to enable users to efficiently understand content even in their busy lives. This system combines a function that automatically generates a summary of the content and provides it in audio format with an emotion engine that recognizes the user's emotions.

[0724] System Configuration

[0725] This system consists of a server, terminals, and an emotion engine. Users can input content information using the terminals and receive summarized content in audio format. Furthermore, the emotion engine can suggest appropriate content and adjust summaries based on the user's emotions.

[0726] Hardware and software to be used

[0727] 1. Servers: Use cloud servers or physical servers with high computing power.

[0728] Database: Use database solutions such as AWS RDS or MongoDB.

[0729] Natural language processing engine: Uses generative AI models such as OpenAI GPT-3.

[0730] Emotion recognition engine: Uses emotion recognition services such as the Microsoft Azure Emotion API.

[0731] Text-to-speech engine: Uses a text-to-speech service such as the Google Text-to-Speech API.

[0732] 2. Device: A smartphone, tablet, or personal computer used by the user.

[0733] User interface: Web browser or dedicated application.

[0734] Program processing flow

[0735] The server receives content information entered by the user, searches the database based on that information, and retrieves the full text of the relevant content. The retrieved full text is then passed to a natural language processing engine, which automatically generates a summary. Furthermore, based on the user's sentiment data obtained from the device, an emotion recognition engine is used to adjust the summary content and voice tone. The adjusted summary is input to a speech synthesis engine and generated as an MP3 audio file. This audio file is stored in temporary storage by the server, and a download link is generated. Finally, the link is sent to the user's device, and the user can play the audio file from this link.

[0736] Specific example

[0737] For example, if a user wants to read a famous novel, they enter the title into the search bar on their device. This information is sent to the server, which retrieves the full text of the novel from its database. The retrieved text is summarized by a generative AI model. Based on this summary, an emotion recognition engine recognizes the user's emotions from their facial expressions and voice tone, and adjusts the content and tone of the summary. The adjusted summary is converted into an MP3 audio file by a speech synthesis engine and provided to the user's device.

[0738] Example of a prompt

[0739] "Please provide an audio summary of Natsume Soseki's 'Kokoro,' with tone adjustments based on emotional recognition."

[0740] "Please automatically summarize the latest business books and generate audio files with tone adjustments based on emotion recognition."

[0741] In this way, the system of the present invention enables users to efficiently understand the key points of content even in their busy lives and to provide personalized information tailored to their emotional state.

[0742] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0743] Step 1: Enter content information

[0744] User: Enter the title or ISBN code of the content you want to read from the terminal's interface. For example, if the user wants to read "a famous novel," they would enter the title in the terminal's search bar.

[0745] Input: Title and ISBN code entered by the user.

[0746] Output: Input data is sent to the server.

[0747] Specific operation: The terminal interface displays "Please enter the title or ISBN of the content." The user enters the information using the keyboard and then clicks the "Search" button.

[0748] Step 2: Search for content data

[0749] Terminal: Sends information about the content entered by the user to the server.

[0750] Server: Based on the received information, it searches a database (e.g., AWS RDS or MongoDB) and finds data with matching content. It then sends the corresponding data back to the terminal in JSON format.

[0751] Input: Entered content information.

[0752] Output: The full text and associated metadata of the relevant content retrieved from the database.

[0753] Specific operation: The terminal sends input data to the server using a REST API. The server executes SQL or NoSQL queries against the database, retrieves the relevant data, converts it to JSON format, and sends it back to the terminal.

[0754] Step 3: Generate an automated summary

[0755] Server: The full text of the content is passed to a natural language processing (NLP) engine (e.g., OpenAI GPT-3 model), which extracts the important parts and generates a summary.

[0756] Input: Full text (JSON format).

[0757] Output: Summarized text.

[0758] Specific operation: The server inputs the full text of the retrieved content into the AI ​​model along with the prompt, "Please summarize:". The AI ​​model analyzes the main parts of the text, extracts the important parts, and generates a summary. The generated summary is saved as a text file.

[0759] Step 4: Adjusting Emotion Recognition and Summarization

[0760] User: While generating a summary, the device's built-in camera and microphone are used to collect user emotion data. For example, when the user is smiling while looking at the screen.

[0761] Server: Uses an emotion engine (e.g., Microsoft Azure Emotion API) that analyzes the user's facial expressions, voice tone, input text, etc., to recognize the user's current emotional state.

[0762] Input: User's facial expression, voice tone, and input text.

[0763] Output: Recognized emotion data.

[0764] Specific operation: The user's device captures facial expressions with its camera and collects audio with its microphone. This data is sent to a server, which uses an emotion engine to send a prompt message saying, "Please recognize the emotion." Based on the recognized emotion data, the content and tone of the summary are adjusted.

[0765] Step 5: Generate audio files

[0766] Server: Inputs the refined summary into a text-to-speech engine (e.g., Google Text-to-Speech API) and converts it into audio data.

[0767] Input: Adjusted summary text.

[0768] Output: MP3 audio file.

[0769] Specific operation: The server inputs the adjusted summary into the speech synthesis engine, along with the prompt, "Please convert this text into speech, including emotional tone." The generated audio data is saved to the server as an MP3 file.

[0770] Step 6: Audio Distribution

[0771] Server: Saves the generated audio files to temporary storage (e.g., Amazon S3) and generates a download link.

[0772] Terminal: Presents the received download link to the user.

[0773] Input: MP3 audio file.

[0774] Output: Download link.

[0775] Specific operation: The server uploads the audio file to temporary storage and generates a public download link. This link is sent to the device, and the device notifies the user of the link.

[0776] Step 7: Play back audio

[0777] User: Select an audio playback option from the device's notifications or interface. For example, the user presses the "Play" button.

[0778] Device: Download the audio file from the download link and play the audio using the built-in player.

[0779] Input: Download link.

[0780] Output: Audio played on the device.

[0781] Specific operation: When the user presses the "Play" button, the device retrieves the audio file from the download link, launches the media player, and plays the audio.

[0782] Step 8: Add to favorites and receive notifications

[0783] User: Adds content they want to read to their favorites list. For example, the user presses the "Add to Favorites" button.

[0784] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[0785] Input: Information about the content to add to your favorites list.

[0786] Output: Notification when a new summary is generated.

[0787] Specific operation: When a user presses the "Add to Favorites" button, information about that content is sent to the server and stored in the database. The server sends an email or app notification to the user when a new summary is added.

[0788] (Application Example 2)

[0789] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0790] In modern society, users lead busy lives and find it difficult to dedicate sufficient time to reading books. Therefore, there is a need for a system that allows for efficient comprehension of book content. However, existing systems merely provide summaries and lack personalization based on user emotions, resulting in low satisfaction. Furthermore, the audio summaries do not adapt to the user's state of mind, leading to insufficient information comprehension and engagement.

[0791] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting information of a book that the user wants to read, means for receiving the book information input by the user and searching for the book in a database, means for summarizing the full text of the searched book, means for converting the summarized book content into an audio file, means for distributing the audio file to the user's terminal, means for recognizing the user's emotions, and means for adjusting the summary content and tone based on the recognized emotions. As a result, users can efficiently understand the content of books even in their busy lives, and receive personalized book summaries based on their emotions in audio format, enabling a more satisfying information provision.

[0792] "User" refers to a person who uses the system to obtain summary information about books.

[0793] "Book information" refers to data used to identify a book, such as the book's title, ISBN code, and author's name.

[0794] A "database" refers to a collection of information that stores the full text of a book and related metadata.

[0795] "Full text" refers to the entire body of text in a book.

[0796] A "summary" refers to information that has been simplified by extracting the most important parts from the full text of a book.

[0797] An "audio file" refers to a digital file in which the content of a summarized book has been converted into audio data.

[0798] "Device" refers to electronic devices used by users, such as smartphones, smart glasses, and head-mounted displays.

[0799] "Emotion recognition" refers to technology that detects and analyzes the emotional state of a user.

[0800] "Tone" refers to the quality or tone of voice, and means a characteristic of voice that is altered based on emotion.

[0801] "Distribution" refers to the act of sending data from a server to a terminal.

[0802] "Personalization" refers to customization tailored to the user's specific needs and preferences.

[0803] Modes for carrying out the invention

[0804] System Overview

[0805] The system of this invention consists of a server, a terminal, and an emotion engine that recognizes the user's emotions. The user can use the terminal to input information about a specific book and obtain a summary of that book as an audio file. Furthermore, the emotion engine is used to adjust the summary according to the user's emotions.

[0806] Hardware and software used

[0807] Server: AWS (Amazon Web Services) or GCP (Google Cloud Platform)

[0808] Database: PostgreSQL or MySQL

[0809] NLP engine: SpaCy or BERT, both Python libraries.

[0810] Emotion Engine: An emotion recognition model using OpenCV and TensorFlow

[0811] Text-to-speech engine: Amazon Polly or Google Text-to-Speech

[0812] Devices: Smartphones, smart glasses, head-mounted displays

[0813] Explanation of the program's processing

[0814] 1. Book Selection

[0815] The user enters the title or ISBN code of the book they want to read through the terminal's interface. For example, if the user wants to read the book "1984," they would enter "1984" into the terminal's search bar.

[0816] 2. Searching for book data

[0817] The server receives the book information entered by the user, searches its database, and identifies the full text of any matching books. It then sends this result back to the terminal.

[0818] 3. Automatic summary generation

[0819] The server uses an NLP engine (e.g., spaCy or BERT) to extract key parts from the full text of the searched books and generate summaries.

[0820] 4. Emotion recognition

[0821] The device uses its camera and microphone to capture the user's facial expressions and voice tone, which are then analyzed by an emotion engine (e.g., OpenCV and TensorFlow models).

[0822] 5. Summarizing and adjusting the tone

[0823] Based on the analyzed sentiment data, the server adjusts the summary content and voice tone to personalize each user. This tone adjustment is performed during the voice file generation process.

[0824] 6. Generating audio files

[0825] The server uses a text-to-speech engine (e.g., Amazon Polly or Google Text-to-Speech) to convert the refined summary into audio data.

[0826] 7. Audio distribution

[0827] The server uploads the generated audio files to cloud storage and provides a download link to the user's device. The user can then retrieve the audio files via this link and make them available for use.

[0828] Specific example

[0829] The user opens the application and types "1984". The camera captures the user's face and determines that they are relaxed. The server uses this emotional information to generate a summary of the book "1984" in a friendly tone and provides it as an audio file. This entire process allows the user to efficiently understand the book's content even in a busy environment.

[0830] Example of a prompt

[0831] "Summarize the key information from all the texts from 1984 in 50 characters or less."

[0832] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0833] Step 1:

[0834] The user enters the title and ISBN code of the book they want to read through the terminal's interface. For example, if the user wants to read "1984," they enter "1984" into the terminal's search bar. This action results in data transfer, where the terminal sends the entered book information to the server.

[0835] Step 2:

[0836] The server queries the database for book information received from the terminal and searches for matching book data (full text and metadata). During this process, the server uses the title and ISBN code as keys to execute the search query. It retrieves the full text of the relevant books as a search result and sends it back to the terminal.

[0837] Step 3:

[0838] The server passes the full text of the book, retrieved from the database, to a natural language processing (NLP) engine. The NLP engine analyzes the text, extracts key parts, and generates a summary. For example, the NLP model might create a summary of the book's main plot and characters in 50 characters or less. This summary is temporarily stored.

[0839] Step 4:

[0840] The device uses its built-in camera and microphone to capture the user's facial expressions and voice tone. The collected data is sent to a server in real time. For example, the camera can capture the user's smile and send it as input data to an emotion recognition model.

[0841] Step 5:

[0842] The server uses an emotion recognition model (e.g., OpenCV and TensorFlow) to analyze the user's emotional state. It recognizes the user's emotions (e.g., joy, sadness, surprise, etc.) from the video and audio provided as input data. This analysis result is used to refine the summary data as an emotional state.

[0843] Step 6:

[0844] The server adjusts the content and tone of the summary generated by the NLP engine based on the user's emotional state. For example, if the user's emotion is "relaxed," the summary's tone will be adjusted to softer language to make it more approachable. As a result of this step, the adjusted summary data is generated.

[0845] Step 7:

[0846] The server inputs the adjusted summary data into a text-to-speech engine (e.g., Amazon Polly or Google Text-to-Speech) to generate an audio file. The text-to-speech engine generates audio data that matches the summary content and emotional tone, and converts it into an MP3 audio file. This audio file is stored on the server.

[0847] Step 8:

[0848] The server uploads the generated audio file to cloud storage and generates a download link. This link is resent to the device, and the user receives a notification. By clicking the link, the user can download and listen to the audio file.

[0849] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0850] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0851] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0852] [Third Embodiment]

[0853] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0854] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0855] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0856] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0857] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0858] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0859] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0860] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0861] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0862] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0863] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0864] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0865] ---

[0866] This invention provides a system that enables users to efficiently understand the content of books even in their busy lives. This system has the function of automatically generating a book summary and providing it in audio format. The program processing of this system is described in detail below.

[0867] System Overview

[0868] This system consists of a server and terminals, allowing users to input book information via the terminals and receive summarized content in audio format. The system has the following main functions:

[0869] 1. Book Selection

[0870] User: Enter the title or ISBN code of the book you want to read using the terminal's interface. For example, if the user wants to read "a famous novel," they would enter the title in the terminal's search bar.

[0871] 2. Searching for book data

[0872] Terminal: Sends the book information entered by the user to the server.

[0873] Server: Based on the received information, it searches the database and finds data for matching books. It then sends the found data back to the terminal. For example, it might send the full text and related metadata of a famous novel to the terminal.

[0874] 3. Automatic summary generation

[0875] Server: The full text of a book is passed to a natural language processing (NLP) engine, which extracts the important parts and generates a short summary. For example, the NLP engine extracts the key plot points and characters from the full text of a famous novel and summarizes them concisely.

[0876] 4. Generating audio files

[0877] Server: Inputs the generated summary into the speech synthesis engine and converts it into audio data. For example, the speech synthesis engine converts the summary content into an MP3 audio file.

[0878] 5. Audio distribution

[0879] Server: Saves the audio file to temporary storage and generates a download link.

[0880] Device: Presents the received download link to the user. For example, a notification might appear stating, "You can listen to an audio summary of a famous novel."

[0881] 6. Audio Playback

[0882] User: Select an audio playback option from the device's notifications or interface. For example, the user presses the "Play" button.

[0883] Device: Download the audio file from the download link and play it using the built-in player.

[0884] 7. Add to favorites and receive notifications

[0885] User: Adds books they want to read to their favorites list. For example, the user presses the "Add to Favorites" button.

[0886] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[0887] In this way, the system of the present invention provides users with a means to efficiently understand the content of books even in their busy lives. By automating each processing step, users can easily access and use the system. Although specific examples have been provided, the system is not limited to these and can be applied in a variety of ways.

[0888] The following describes the processing flow.

[0889] ---

[0890] Step 1:

[0891] User: Open the dedicated application or webpage on your device and enter the title or ISBN code of the book you want to read.

[0892] Step 2:

[0893] Terminal: Sends the book information entered by the user to the server.

[0894] Step 3:

[0895] Server: Based on the title or ISBN code of the received book, the server searches the database for data on matching books.

[0896] Step 4:

[0897] Server: Returns the full text of the searched book and related metadata to the terminal.

[0898] Step 5:

[0899] Server: Passes the full text to a natural language processing (NLP) engine, which extracts the important parts and generates a summary.

[0900] Step 6:

[0901] Server: Inputs the generated summary into the speech synthesis engine and converts it into audio data. Specifically, the speech synthesis engine converts the summary content into an MP3 audio file.

[0902] Step 7:

[0903] Server: Saves the generated audio files to temporary storage and generates a download link.

[0904] Step 8:

[0905] Server: Sends a download link to the device.

[0906] Step 9:

[0907] Terminal: Receives the sent download link and displays a notification to the user stating, "The summary is available for audio playback."

[0908] Step 10:

[0909] User: Select the audio playback option in the device's notification or interface and press the "Play" button.

[0910] Step 11:

[0911] Device: Download the audio file from the download link and play the audio using the built-in player.

[0912] Step 12:

[0913] User: Add books you want to read to your favorites list as needed. For example, the user presses the "Add to Favorites" button.

[0914] Step 13:

[0915] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[0916] ---

[0917] The above is a detailed explanation of the processing steps. This system automates the entire process, from inputting book information to audio playback, to help users efficiently understand the content of books.

[0918] (Example 1)

[0919] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0920] In today's busy lives, many people don't have the time to efficiently understand the content of books. There is a growing demand for systems that allow users to easily access summarized book content and listen to it as audio. Traditional systems lacked convenience because the processes of generating summaries, converting to audio files, and distribution were time-consuming and laborious.

[0921] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0922] In this invention, the server includes means for inputting information about a book that the user wants to read, means for receiving the book information input by the user and searching for the book in a data store, means for summarizing the entire text of the searched book, means for inputting the generated summary into a natural language processing engine, means for inputting the generated summary into a speech synthesis engine and converting it into audio data, means for temporarily storing the audio data and generating a download link thereof, and means for delivering the download link to the user's terminal. This makes it possible for the user to efficiently and quickly obtain an audio summary of the book.

[0923] A "user" is an individual or organization that uses the system to obtain book summaries and audio data.

[0924] "Book information" refers to information necessary to uniquely identify a book, such as the book's title, author's name, and ISBN code.

[0925] A "data store" is a storage device or database system used to store and manage book data.

[0926] "The full text of the book" refers to all the text and data contained in the book.

[0927] A "summary" refers to a concise compilation of the most important parts and key points extracted from the entire text of a book.

[0928] A "natural language processing engine" is software or algorithms that analyze text data and perform tasks such as summarizing, translating, and other natural language processing.

[0929] A "speech synthesis engine" is a software or hardware system that takes text data as input and converts it into speech data.

[0930] "Audio data" refers to audio files in an audible format, generated by a speech synthesis engine.

[0931] "Temporary storage" refers to the process of saving data only for a certain period, with deletion or updating performed as needed.

[0932] A "download link" refers to a URL or path used to obtain audio data or other data.

[0933] A "terminal" is an electronic device used by a user to access and operate a system.

[0934] This invention provides a system that enables users to efficiently understand the content of books even in their busy lives. This system has the function of automatically generating a book summary and providing it in audio format. Specific embodiments of this system are described in detail below.

[0935] System Configuration

[0936] This system primarily consists of a server and terminals, and users obtain book summaries as audio via the terminals. The system is comprised of the following hardware and software combinations.

[0937] hardware

[0938] The server includes a data store (e.g., a MySQL database), an NLP engine, and a speech synthesis engine.

[0939] Device: An electronic device used by a user, such as a smartphone, tablet, or personal computer.

[0940] software

[0941] Natural language processing engine (e.g., GPT-3)

[0942] Text-to-speech engine (e.g., Google Text-to-Speech)

[0943] Database management system (e.g., MySQL)

[0944] HTTP request processing libraries (e.g., Axios)

[0945] Operation flow

[0946] The operation flow of this system is as follows:

[0947] 1. Book Selection

[0948] The user enters the title or ISBN code of the book they want to read through the terminal's interface. For example, if the user wants to read "a famous novel," they enter the title in the terminal's search bar and press the "Search" button.

[0949] 2. Searching for book data

[0950] The terminal sends the book information entered by the user to the server. Based on the received information, the server searches its data store for matching book data. As a result, it sends back the full text of "a famous novel" and related metadata to the terminal.

[0951] 3. Automatic summary generation

[0952] The server passes the full text of the book to a natural language processing engine, which extracts the important parts and generates a short summary. For example, it passes a prompt sentence like the following to the natural language processing engine:

[0953] Please summarize the following text: "The complete text of a famous novel."

[0954] 4. Generating audio files

[0955] The server inputs the generated summary into the speech synthesis engine and converts it into speech data. Specifically, it uses prompts like the following:

[0956] "Please convert this summary into an audio file."

[0957] The generated audio file will be temporarily saved.

[0958] 5. Audio distribution

[0959] The server generates a download link for the audio file and notifies the terminal of that link. For example, it sends the user a message saying, "You can listen to an audio summary of a famous novel." The terminal then displays the received download link on its interface.

[0960] 6. Audio Playback

[0961] The user selects the audio playback option through the device's notifications or interface. When the user presses the "Play" button, the device retrieves the audio file from the download link and plays it using the built-in player.

[0962] 7. Add to favorites and receive notifications

[0963] Users can add books they want to read to their favorites list. When a new summary is generated, the server notifies the user via email or app notification.

[0964] This system enables users to efficiently understand the content of books even in their busy lives, and by automating each processing step, it achieves simple access and use. For example, it can be used in the same process when a user is trying to read the latest business book, "The Road to Success."

[0965] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0966] Step 1:

[0967] User: Enter information about the book you want to read. Specifically, the user enters the book title and ISBN code through the terminal interface and presses the "Search" button. This entered information is then sent to the next processing step.

[0968] Step 2:

[0969] Terminal: Sends the book information entered by the user to the server. Specifically, it sends the book information to the server using an HTTP request. This operation passes the input data to the server.

[0970] Step 3:

[0971] Server: Searches the data store based on the received book information. Specifically, it queries the data store (e.g., a MySQL database) based on the received book information and retrieves the full text and metadata of matching books. This query process outputs the full text and metadata.

[0972] Step 4:

[0973] Server: The server passes the full text obtained as search results to a natural language processing (NLP) engine to generate a summary. Specifically, it inputs the full text into the NLP engine using the prompt message, "Summarize the following text: The full text of 'Book Title'." The NLP engine performs text analysis, extracts the important parts, and generates a summary. This summary is then passed to the next step.

[0974] Step 5:

[0975] Server: Inputs the generated summary into a text-to-speech engine and converts it into audio data. Specifically, it inputs the summary into a text-to-speech engine (e.g., Google Text-to-Speech) using a prompt message such as "Please convert this summary into an audio file." The text-to-speech engine converts the text into speech and generates an audio file in MP3 format. This audio file is then passed on to the next step.

[0976] Step 6:

[0977] Server: Temporarily stores the audio file and generates a download link. Specifically, it uploads the generated audio file to cloud storage and generates a download link. This link is then sent to the next step.

[0978] Step 7:

[0979] Device: Notifies the user of the generated download link. Specifically, the device's interface and notification system display the link along with the message, "You can listen to an audio summary of the book title." When the user clicks this link, they proceed to the next step.

[0980] Step 8:

[0981] User: Selects the option to play the audio file. Specifically, the user clicks the received link and plays it using the device's built-in player. This playback action allows the user to listen to the summarized content in audio.

[0982] Step 9:

[0983] User: Adds books they want to read to their favorites list. Specifically, they press the "Add to Favorites" button on their device. This action saves the book's information to the user's personal data store.

[0984] Step 10:

[0985] Server: When a new summary is generated for a book added to the favorites list, a notification is sent. Specifically, when new summary data is generated, the user is notified via email or app notification. This notification allows the user to check the new information immediately.

[0986] (Application Example 1)

[0987] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0988] For users leading busy lives, there is a need for efficient ways to understand the content of books. However, traditional methods of providing book summaries and audiobooks sometimes make it difficult for users to quickly obtain the information they are looking for. Furthermore, there is a lack of systems that easily recommend new reading material and provide easy access to desired reading. A new system is needed to address these challenges.

[0989] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0990] In this invention, the server includes means for the user to input information about a reading material, means for receiving the information about the reading material input by the user and searching for the reading material in a database, means for summarizing the entire text of the searched reading material, means for converting the summarized content of the reading material into an audio file, means for distributing the audio file to the user's terminal device, and means for presenting recommended reading material based on the information about the reading material input by the user. As a result, the user can not only quickly listen to a summary in audio format based on the information they have input, but also receive recommendations for related new reading material, enabling them to understand the content of books more easily and efficiently.

[0991] A "user" is an individual or organization that uses the system to input information about reading material and receives the retrieved summaries and audio files.

[0992] "Reading materials" is a general term for information media provided in text format, such as books, magazines, articles, and ebooks.

[0993] "Means of input" refers to methods of providing information to a system using keyboards, voice input, touchscreens, etc.

[0994] "Means of receiving" refers to the function by which a server retrieves information sent by a user.

[0995] "Searching methods" refer to the function of finding relevant reading material from a database based on the received information.

[0996] "Full text" refers to text data that corresponds to the entire text of the searched reading material.

[0997] "Methods of summarization" refer to the function of extracting important parts from the entire text and summarizing them concisely.

[0998] "Means of converting to audio files" refers to a function that converts summarized text into audio data.

[0999] "Terminal devices" refer to electronic devices such as smartphones, tablets, and computers.

[1000] "Means of distribution" refers to the function of sending the generated audio file to the user's terminal device.

[1001] "Means for suggesting recommended reading material" refers to a function that suggests new reading material related to the user's input information.

[1002] The "playback option" refers to the choices a user can make to listen to an audio file.

[1003] "Adding to a favorites list" refers to a function that allows users to save reading material to a list so they can easily access it later.

[1004] "Means of sending notifications" refers to a function that allows the server to notify users when a new summary is generated.

[1005] This invention is a system designed to enable users to efficiently understand the content of reading materials. This system consists of a server and terminal devices, allowing users to input book information and obtain summarized content in audio format. A specific embodiment of this system is described below.

[1006] Hardware and software to be used

[1007] Hardware:

[1008] Terminal devices: Smartphones, tablets, computers, smart glasses, head-mounted displays, etc.

[1009] Server: A high-performance computer with internet connectivity.

[1010] software:

[1011] Python programming language

[1012] requests library: Used for communication between servers and terminal devices.

[1013] gTTS (Google Text-to-Speech) library: Converts text to speech.

[1014] playsound library: Plays the generated audio file.

[1015] System operation

[1016] Book Search

[1017] The user inputs the title and ISBN code of the book they want to read through the interface of the terminal device, and sends that information to the server. For example, the user inputs the title of "a famous novel."

[1018] Book summary generation

[1019] The server searches the database based on the input information and retrieves data on matching reading materials. The retrieved complete text is then summarized using a natural language processing engine. This summarization process extracts the main points and important parts of the text and condenses them into a concise form.

[1020] Audio file generation and distribution

[1021] The gTTS library is used to convert the generated summary text into an audio file. The generated audio file is delivered from the server to the terminal device, and the user can understand the summary by listening to it. For example, the content of a summarized "famous novel" is delivered as an audio file.

[1022] Recommended reading material

[1023] Based on the information entered about the reading material, the server generates and presents relevant recommended reading materials to the user. This allows the user to discover new and interesting reading material.

[1024] Specific example

[1025] For example, if a user enters the title "I Am a Cat," the system searches the database for the book, summarizes the entire text, and converts it into an audio file. This audio file is then delivered to the user's terminal device, allowing them to quickly listen to the content of "I Am a Cat." In addition, "Other Works by Soseki" are presented as recommended reading material.

[1026] Example of a prompt

[1027] Please enter the book title: I Am a Cat

[1028] Thus, the system of the present invention is a system that can efficiently provide summaries of reading materials and improve the reading experience by combining natural language processing technology and speech synthesis technology.

[1029] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1030] Step 1:

[1031] The user inputs information about the book they are reading from a terminal device. The input information, such as the book title and ISBN code, is transmitted to the server through the terminal device's interface.

[1032] Input: Book title and ISBN code

[1033] Output: Search information sent to the server

[1034] Step 2:

[1035] The server searches the database based on the information it receives about the reading material. The database stores the complete text of numerous reading materials and associated metadata.

[1036] Input: Search information (book title and ISBN code)

[1037] Output: Data on matching reading materials

[1038] Step 3:

[1039] The server passes the entire text retrieved as search results to a natural language processing (NLP) engine, which then generates a summary. The NLP engine extracts the important parts and creates a compact summary text.

[1040] Input: Entire text

[1041] Output: Summary text

[1042] Step 4:

[1043] The server inputs the generated summary text into a speech synthesis engine (e.g., the gTTS library) and converts it into an audio file. The audio file format commonly used is MP3.

[1044] Input: Summary text

[1045] Output: Audio file (MP3 format)

[1046] Step 5:

[1047] The server saves the generated audio file to temporary storage and generates a download link. The server then sends this link to the terminal device.

[1048] Input: Audio file (MP3 format)

[1049] Output: Download link

[1050] Step 6:

[1051] The user receives a download link from the terminal device, retrieves the audio file, and plays it on the terminal device's player. The terminal device then displays notifications and playback options.

[1052] Input: Download link

[1053] Output: Playback of audio file

[1054] Step 7:

[1055] Users add books they want to read to their favorites list. The server notifies the user via email or app notification when a new summary is generated.

[1056] Input: Registration information for your favorites list

[1057] Output: Notification of new summary

[1058] Step 8:

[1059] The server generates recommended reading material based on the user's input and presents it to the terminal device. Similar genres and authors are considered in the recommendations.

[1060] Input: User input information

[1061] Output: Presentation of recommended reading material

[1062] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1063] ---

[1064] This invention provides a system that enables users to efficiently understand the content of books even in their busy lives. This system combines a function that automatically generates book summaries and provides them in audio format with an emotion engine that recognizes the user's emotions.

[1065] System Overview

[1066] This system consists of a server, terminals, and an emotion engine. Users can input book information using the terminals and receive summarized content in audio. Furthermore, the emotion engine can suggest appropriate books and adjust summaries based on the user's emotions.

[1067] 1. Book Selection

[1068] User: Enter the title or ISBN code of the book you want to read using the terminal's interface. For example, if the user wants to read "a famous novel," they would enter the title in the terminal's search bar.

[1069] 2. Searching for book data

[1070] Terminal: Sends the book information entered by the user to the server.

[1071] Server: Based on the received information, it searches the database and finds data for matching books. It then sends the found data back to the terminal. For example, it might send the full text and related metadata of a famous novel to the terminal.

[1072] 3. Automatic summary generation

[1073] Server: The full text of a book is passed to a natural language processing (NLP) engine, which extracts the important parts and generates a summary. For example, the NLP engine extracts the key plot points and characters from the full text of a famous novel and summarizes them concisely.

[1074] 4. Adjusting emotion recognition and summarization

[1075] Server: Uses an emotion engine that analyzes the user's facial expressions, voice tone, input text, etc., to recognize the user's current emotional state. For example, the emotion engine might detect a happy emotion from the user's voice tone.

[1076] Server: Based on recognized emotions, adjusts the content and tone of the summary and audio file. For example, the emotion engine changes the tone of the summary according to the user's emotions, adjusting it to emphasize positive content.

[1077] 5. Generating audio files

[1078] Server: Inputs the adjusted summary into the speech synthesis engine and converts it into audio data. Specifically, the speech synthesis engine converts the summary into an MP3 audio file with an emotion-appropriate tone.

[1079] 6. Audio distribution

[1080] Server: Saves the generated audio files to temporary storage and generates a download link.

[1081] Device: Presents the received download link to the user. For example, a notification might appear stating, "You can listen to an audio summary of a famous novel."

[1082] 7. Audio Playback

[1083] User: Select an audio playback option from the device's notifications or interface. For example, the user presses the "Play" button.

[1084] Device: Download the audio file from the download link and play the audio using the built-in player.

[1085] 8. Add to favorites and receive notifications

[1086] User: Adds books they want to read to their favorites list. For example, the user presses the "Add to Favorites" button.

[1087] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[1088] In this way, the system of the present invention enables users to efficiently understand the content of books even in their busy lives, and further provides a personalized experience through the emotion engine. Each processing step is automated, making it easy for users to access and use. While specific examples have been provided, the system is not limited to these and can be applied in a variety of ways.

[1089] The following describes the processing flow.

[1090] ---

[1091] Step 1:

[1092] User: Open the dedicated application or webpage on your device and enter the title or ISBN code of the book you want to read. For example, if the user wants to read "a famous novel," they would enter the title in the search bar on their device.

[1093] Step 2:

[1094] Terminal: Sends the book information entered by the user to the server.

[1095] Step 3:

[1096] Server: Based on the title or ISBN code of the received book, the server searches the database for data on matching books.

[1097] Step 4:

[1098] Server: Returns the full text of the searched book and related metadata to the terminal.

[1099] Step 5:

[1100] Server: The full text of a book is passed to a natural language processing (NLP) engine, which extracts the important parts and generates a summary. For example, the NLP engine extracts the key plot points and characters from the full text of a famous novel and summarizes them concisely.

[1101] Step 6:

[1102] Server: After the summary is generated, an emotion engine analyzes the user's facial expressions, voice tone, input text, etc., to recognize the user's emotional state. For example, the emotion engine might detect a cheerful emotion from the user's voice tone.

[1103] Step 7:

[1104] Server: Adjusts the content and tone of the summary and audio file based on recognized emotions. For example, the emotion engine might change the tone of the summary to a positive one based on the user's positive emotions.

[1105] Step 8:

[1106] Server: Inputs the adjusted summary into the speech synthesis engine and converts it into audio data. Specifically, the speech synthesis engine converts the summary into an MP3 audio file with an emotion-appropriate tone.

[1107] Step 9:

[1108] Server: Saves the generated audio files to temporary storage and generates a download link.

[1109] Step 10:

[1110] Server: Sends a download link to the device.

[1111] Step 11:

[1112] Terminal: Receives the sent download link and displays a notification to the user stating, "The summary is available for audio playback."

[1113] Step 12:

[1114] User: Select the audio playback option in the device's notification or interface and press the "Play" button.

[1115] Step 13:

[1116] Device: Download the audio file from the download link and play the audio using the built-in player.

[1117] Step 14:

[1118] User: Add books you want to read to your favorites list as needed. For example, the user presses the "Add to Favorites" button.

[1119] Step 15:

[1120] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[1121] ---

[1122] The above is a detailed explanation of the processing steps. This system automates the entire process from book information input to audio playback, and further provides a personalized experience through its emotion engine. Users can easily access and use it.

[1123] (Example 2)

[1124] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1125] In modern society, users are required to consume information and content efficiently amidst their busy lives. However, traditional methods present challenges such as the lack of time to read long texts like books and the effort required to understand summarized information. Furthermore, providing personalized information tailored to the user's current emotional state is difficult.

[1126] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting information of content that the user wants to read, means for receiving the information of content input by the user and searching for the content from a database, means for automatically summarizing the full text of the searched content, means for adjusting the summarized content based on the user's emotional state, means for converting the adjusted content into an audio file, and means for delivering the audio file to the user's device. As a result, even in a busy life, the user can efficiently understand the main points of the content and receive personalized information tailored to their emotional state.

[1127] "User" refers to an individual or legal entity that intends to use this system to search for, summarize, or convert content into audio.

[1128] "Content" refers to information such as books, articles, and reports, and includes informational materials that users wish to read.

[1129] "Means of inputting information" refers to the interface that allows users to input information such as the content title, author name, and ISBN code into their device.

[1130] A "database" refers to a collection of information that stores, manages, and allows for quick searching of content metadata and full text.

[1131] "Methods for automatic summarization" refer to algorithms or engines that use natural language processing technology to extract key parts from the full text of content and generate a concise summary.

[1132] "Means of adjustment based on emotional state" refers to a function that analyzes the user's current emotions and adjusts the content and tone of information such as summaries and audio according to those emotions.

[1133] "Means of converting to audio files" refers to speech synthesis engines and software that convert automatically summarized text information into audio data.

[1134] "Device" refers to the electronic device that a user uses to access this system, and includes smartphones, tablets, personal computers, etc.

[1135] "Means of distribution" refers to the function of transmitting generated audio files over a network in order to provide them to users.

[1136] This invention is a system designed to enable users to efficiently understand content even in their busy lives. This system combines a function that automatically generates a summary of the content and provides it in audio format with an emotion engine that recognizes the user's emotions.

[1137] System Configuration

[1138] This system consists of a server, terminals, and an emotion engine. Users can input content information using the terminals and receive summarized content in audio format. Furthermore, the emotion engine can suggest appropriate content and adjust summaries based on the user's emotions.

[1139] Hardware and software to be used

[1140] 1. Servers: Use cloud servers or physical servers with high computing power.

[1141] Database: Use database solutions such as AWS RDS or MongoDB.

[1142] Natural language processing engine: Uses generative AI models such as OpenAI GPT-3.

[1143] Emotion recognition engine: Uses emotion recognition services such as the Microsoft Azure Emotion API.

[1144] Text-to-speech engine: Uses a text-to-speech service such as the Google Text-to-Speech API.

[1145] 2. Device: A smartphone, tablet, or personal computer used by the user.

[1146] User interface: Web browser or dedicated application.

[1147] Program processing flow

[1148] The server receives content information entered by the user, searches the database based on that information, and retrieves the full text of the relevant content. The retrieved full text is then passed to a natural language processing engine, which automatically generates a summary. Furthermore, based on the user's sentiment data obtained from the device, an emotion recognition engine is used to adjust the summary content and voice tone. The adjusted summary is input to a speech synthesis engine and generated as an MP3 audio file. This audio file is stored in temporary storage by the server, and a download link is generated. Finally, the link is sent to the user's device, and the user can play the audio file from this link.

[1149] Specific example

[1150] For example, if a user wants to read a famous novel, they enter the title into the search bar on their device. This information is sent to the server, which retrieves the full text of the novel from its database. The retrieved text is summarized by a generative AI model. Based on this summary, an emotion recognition engine recognizes the user's emotions from their facial expressions and voice tone, and adjusts the content and tone of the summary. The adjusted summary is converted into an MP3 audio file by a speech synthesis engine and provided to the user's device.

[1151] Example of a prompt

[1152] "Please provide an audio summary of Natsume Soseki's 'Kokoro,' with tone adjustments based on emotional recognition."

[1153] "Please automatically summarize the latest business books and generate audio files with tone adjustments based on emotion recognition."

[1154] In this way, the system of the present invention enables users to efficiently understand the key points of content even in their busy lives and to provide personalized information tailored to their emotional state.

[1155] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1156] Step 1: Enter content information

[1157] User: Enter the title or ISBN code of the content you want to read from the terminal's interface. For example, if the user wants to read "a famous novel," they would enter the title in the terminal's search bar.

[1158] Input: Title and ISBN code entered by the user.

[1159] Output: Input data is sent to the server.

[1160] Specific operation: The terminal interface displays "Please enter the title or ISBN of the content." The user enters the information using the keyboard and then clicks the "Search" button.

[1161] Step 2: Search for content data

[1162] Terminal: Sends information about the content entered by the user to the server.

[1163] Server: Based on the received information, it searches a database (e.g., AWS RDS or MongoDB) and finds data with matching content. It then sends the corresponding data back to the terminal in JSON format.

[1164] Input: Entered content information.

[1165] Output: The full text and associated metadata of the relevant content retrieved from the database.

[1166] Specific operation: The terminal sends input data to the server using a REST API. The server executes SQL or NoSQL queries against the database, retrieves the relevant data, converts it to JSON format, and sends it back to the terminal.

[1167] Step 3: Generate an automated summary

[1168] Server: The full text of the content is passed to a natural language processing (NLP) engine (e.g., OpenAI GPT-3 model), which extracts the important parts and generates a summary.

[1169] Input: Full text (JSON format).

[1170] Output: Summarized text.

[1171] Specific operation: The server inputs the full text of the retrieved content into the AI ​​model along with the prompt, "Please summarize:". The AI ​​model analyzes the main parts of the text, extracts the important parts, and generates a summary. The generated summary is saved as a text file.

[1172] Step 4: Adjusting Emotion Recognition and Summarization

[1173] User: While generating a summary, the device's built-in camera and microphone are used to collect user emotion data. For example, when the user is smiling while looking at the screen.

[1174] Server: Uses an emotion engine (e.g., Microsoft Azure Emotion API) that analyzes the user's facial expressions, voice tone, input text, etc., to recognize the user's current emotional state.

[1175] Input: User's facial expression, voice tone, and input text.

[1176] Output: Recognized emotion data.

[1177] Specific operation: The user's device captures facial expressions with its camera and collects audio with its microphone. This data is sent to a server, which uses an emotion engine to send a prompt message saying, "Please recognize the emotion." Based on the recognized emotion data, the content and tone of the summary are adjusted.

[1178] Step 5: Generate audio files

[1179] Server: Inputs the refined summary into a text-to-speech engine (e.g., Google Text-to-Speech API) and converts it into audio data.

[1180] Input: Adjusted summary text.

[1181] Output: MP3 audio file.

[1182] Specific operation: The server inputs the adjusted summary into the speech synthesis engine, along with the prompt, "Please convert this text into speech, including emotional tone." The generated audio data is saved to the server as an MP3 file.

[1183] Step 6: Audio Distribution

[1184] Server: Saves the generated audio files to temporary storage (e.g., Amazon S3) and generates a download link.

[1185] Terminal: Presents the received download link to the user.

[1186] Input: MP3 audio file.

[1187] Output: Download link.

[1188] Specific operation: The server uploads the audio file to temporary storage and generates a public download link. This link is sent to the device, and the device notifies the user of the link.

[1189] Step 7: Play back audio

[1190] User: Select an audio playback option from the device's notifications or interface. For example, the user presses the "Play" button.

[1191] Device: Download the audio file from the download link and play the audio using the built-in player.

[1192] Input: Download link.

[1193] Output: Audio played on the device.

[1194] Specific operation: When the user presses the "Play" button, the device retrieves the audio file from the download link, launches the media player, and plays the audio.

[1195] Step 8: Add to favorites and receive notifications

[1196] User: Adds content they want to read to their favorites list. For example, the user presses the "Add to Favorites" button.

[1197] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[1198] Input: Information about the content to add to your favorites list.

[1199] Output: Notification when a new summary is generated.

[1200] Specific operation: When a user presses the "Add to Favorites" button, information about that content is sent to the server and stored in the database. The server sends an email or app notification to the user when a new summary is added.

[1201] (Application Example 2)

[1202] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1203] In modern society, users lead busy lives and find it difficult to dedicate sufficient time to reading books. Therefore, there is a need for a system that allows for efficient comprehension of book content. However, existing systems merely provide summaries and lack personalization based on user emotions, resulting in low satisfaction. Furthermore, the audio summaries do not adapt to the user's state of mind, leading to insufficient information comprehension and engagement.

[1204] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting information of a book that the user wants to read, means for receiving the book information input by the user and searching for the book in a database, means for summarizing the full text of the searched book, means for converting the summarized book content into an audio file, means for distributing the audio file to the user's terminal, means for recognizing the user's emotions, and means for adjusting the summary content and tone based on the recognized emotions. As a result, users can efficiently understand the content of books even in their busy lives, and receive personalized book summaries based on their emotions in audio format, enabling a more satisfying information provision.

[1205] "User" refers to a person who uses the system to obtain summary information about books.

[1206] "Book information" refers to data used to identify a book, such as the book's title, ISBN code, and author's name.

[1207] A "database" refers to a collection of information that stores the full text of a book and related metadata.

[1208] "Full text" refers to the entire body of text in a book.

[1209] A "summary" refers to information that has been simplified by extracting the most important parts from the full text of a book.

[1210] An "audio file" refers to a digital file in which the content of a summarized book has been converted into audio data.

[1211] "Device" refers to electronic devices used by users, such as smartphones, smart glasses, and head-mounted displays.

[1212] "Emotion recognition" refers to technology that detects and analyzes the emotional state of a user.

[1213] "Tone" refers to the quality or tone of voice, and means a characteristic of voice that is altered based on emotion.

[1214] "Distribution" refers to the act of sending data from a server to a terminal.

[1215] "Personalization" refers to customization tailored to the user's specific needs and preferences.

[1216] Modes for carrying out the invention

[1217] System Overview

[1218] The system of this invention consists of a server, a terminal, and an emotion engine that recognizes the user's emotions. The user can use the terminal to input information about a specific book and obtain a summary of that book as an audio file. Furthermore, the emotion engine is used to adjust the summary according to the user's emotions.

[1219] Hardware and software used

[1220] Server: AWS (Amazon Web Services) or GCP (Google Cloud Platform)

[1221] Database: PostgreSQL or MySQL

[1222] NLP engine: SpaCy or BERT, both Python libraries.

[1223] Emotion Engine: An emotion recognition model using OpenCV and TensorFlow

[1224] Text-to-speech engine: Amazon Polly or Google Text-to-Speech

[1225] Devices: Smartphones, smart glasses, head-mounted displays

[1226] Explanation of the program's processing

[1227] 1. Book Selection

[1228] The user enters the title or ISBN code of the book they want to read through the terminal's interface. For example, if the user wants to read the book "1984," they would enter "1984" into the terminal's search bar.

[1229] 2. Searching for book data

[1230] The server receives the book information entered by the user, searches its database, and identifies the full text of any matching books. It then sends this result back to the terminal.

[1231] 3. Automatic summary generation

[1232] The server uses an NLP engine (e.g., spaCy or BERT) to extract key parts from the full text of the searched books and generate summaries.

[1233] 4. Emotion recognition

[1234] The device uses its camera and microphone to capture the user's facial expressions and voice tone, which are then analyzed by an emotion engine (e.g., OpenCV and TensorFlow models).

[1235] 5. Summarizing and adjusting the tone

[1236] Based on the analyzed sentiment data, the server adjusts the summary content and voice tone to personalize each user. This tone adjustment is performed during the voice file generation process.

[1237] 6. Generating audio files

[1238] The server uses a text-to-speech engine (e.g., Amazon Polly or Google Text-to-Speech) to convert the refined summary into audio data.

[1239] 7. Audio distribution

[1240] The server uploads the generated audio files to cloud storage and provides a download link to the user's device. The user can then retrieve the audio files via this link and make them available for use.

[1241] Specific example

[1242] The user opens the application and types "1984". The camera captures the user's face and determines that they are relaxed. The server uses this emotional information to generate a summary of the book "1984" in a friendly tone and provides it as an audio file. This entire process allows the user to efficiently understand the book's content even in a busy environment.

[1243] Example of a prompt

[1244] "Summarize the key information from all the texts from 1984 in 50 characters or less."

[1245] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1246] Step 1:

[1247] The user enters the title and ISBN code of the book they want to read through the terminal's interface. For example, if the user wants to read "1984," they enter "1984" into the terminal's search bar. This action results in data transfer, where the terminal sends the entered book information to the server.

[1248] Step 2:

[1249] The server queries the database for book information received from the terminal and searches for matching book data (full text and metadata). During this process, the server uses the title and ISBN code as keys to execute the search query. It retrieves the full text of the relevant books as a search result and sends it back to the terminal.

[1250] Step 3:

[1251] The server passes the full text of the book, retrieved from the database, to a natural language processing (NLP) engine. The NLP engine analyzes the text, extracts key parts, and generates a summary. For example, the NLP model might create a summary of the book's main plot and characters in 50 characters or less. This summary is temporarily stored.

[1252] Step 4:

[1253] The device uses its built-in camera and microphone to capture the user's facial expressions and voice tone. The collected data is sent to a server in real time. For example, the camera can capture the user's smile and send it as input data to an emotion recognition model.

[1254] Step 5:

[1255] The server uses an emotion recognition model (e.g., OpenCV and TensorFlow) to analyze the user's emotional state. It recognizes the user's emotions (e.g., joy, sadness, surprise, etc.) from the video and audio provided as input data. This analysis result is used to refine the summary data as an emotional state.

[1256] Step 6:

[1257] The server adjusts the content and tone of the summary generated by the NLP engine based on the user's emotional state. For example, if the user's emotion is "relaxed," the summary's tone will be adjusted to softer language to make it more approachable. As a result of this step, the adjusted summary data is generated.

[1258] Step 7:

[1259] The server inputs the adjusted summary data into a text-to-speech engine (e.g., Amazon Polly or Google Text-to-Speech) to generate an audio file. The text-to-speech engine generates audio data that matches the summary content and emotional tone, and converts it into an MP3 audio file. This audio file is stored on the server.

[1260] Step 8:

[1261] The server uploads the generated audio file to cloud storage and generates a download link. This link is resent to the device, and the user receives a notification. By clicking the link, the user can download and listen to the audio file.

[1262] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1263] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1264] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1265] [Fourth Embodiment]

[1266] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1267] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1268] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1269] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1270] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1271] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1272] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1273] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1274] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1275] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1276] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1277] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1278] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1279] ---

[1280] This invention provides a system that enables users to efficiently understand the content of books even in their busy lives. This system has the function of automatically generating a book summary and providing it in audio format. The program processing of this system is described in detail below.

[1281] System Overview

[1282] This system consists of a server and terminals, allowing users to input book information via the terminals and receive summarized content in audio format. The system has the following main functions:

[1283] 1. Book Selection

[1284] User: Enter the title or ISBN code of the book you want to read using the terminal's interface. For example, if the user wants to read "a famous novel," they would enter the title in the terminal's search bar.

[1285] 2. Searching for book data

[1286] Terminal: Sends the book information entered by the user to the server.

[1287] Server: Based on the received information, it searches the database and finds data for matching books. It then sends the found data back to the terminal. For example, it might send the full text and related metadata of a famous novel to the terminal.

[1288] 3. Automatic summary generation

[1289] Server: The full text of a book is passed to a natural language processing (NLP) engine, which extracts the important parts and generates a short summary. For example, the NLP engine extracts the key plot points and characters from the full text of a famous novel and summarizes them concisely.

[1290] 4. Generating audio files

[1291] Server: Inputs the generated summary into the speech synthesis engine and converts it into audio data. For example, the speech synthesis engine converts the summary content into an MP3 audio file.

[1292] 5. Audio distribution

[1293] Server: Saves the audio file to temporary storage and generates a download link.

[1294] Device: Presents the received download link to the user. For example, a notification might appear stating, "You can listen to an audio summary of a famous novel."

[1295] 6. Audio Playback

[1296] User: Select an audio playback option from the device's notifications or interface. For example, the user presses the "Play" button.

[1297] Device: Download the audio file from the download link and play it using the built-in player.

[1298] 7. Add to favorites and receive notifications

[1299] User: Adds books they want to read to their favorites list. For example, the user presses the "Add to Favorites" button.

[1300] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[1301] In this way, the system of the present invention provides users with a means to efficiently understand the content of books even in their busy lives. By automating each processing step, users can easily access and use the system. Although specific examples have been provided, the system is not limited to these and can be applied in a variety of ways.

[1302] The following describes the processing flow.

[1303] ---

[1304] Step 1:

[1305] User: Open the dedicated application or webpage on your device and enter the title or ISBN code of the book you want to read.

[1306] Step 2:

[1307] Terminal: Sends the book information entered by the user to the server.

[1308] Step 3:

[1309] Server: Based on the title or ISBN code of the received book, the server searches the database for data on matching books.

[1310] Step 4:

[1311] Server: Returns the full text of the searched book and related metadata to the terminal.

[1312] Step 5:

[1313] Server: Passes the full text to a natural language processing (NLP) engine, which extracts the important parts and generates a summary.

[1314] Step 6:

[1315] Server: Inputs the generated summary into the speech synthesis engine and converts it into audio data. Specifically, the speech synthesis engine converts the summary content into an MP3 audio file.

[1316] Step 7:

[1317] Server: Saves the generated audio files to temporary storage and generates a download link.

[1318] Step 8:

[1319] Server: Sends a download link to the device.

[1320] Step 9:

[1321] Terminal: Receives the sent download link and displays a notification to the user stating, "The summary is available for audio playback."

[1322] Step 10:

[1323] User: Select the audio playback option in the device's notification or interface and press the "Play" button.

[1324] Step 11:

[1325] Device: Download the audio file from the download link and play the audio using the built-in player.

[1326] Step 12:

[1327] User: Add books you want to read to your favorites list as needed. For example, the user presses the "Add to Favorites" button.

[1328] Step 13:

[1329] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[1330] ---

[1331] The above is a detailed explanation of the processing steps. This system automates the entire process, from inputting book information to audio playback, to help users efficiently understand the content of books.

[1332] (Example 1)

[1333] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1334] In today's busy lives, many people don't have the time to efficiently understand the content of books. There is a growing demand for systems that allow users to easily access summarized book content and listen to it as audio. Traditional systems lacked convenience because the processes of generating summaries, converting to audio files, and distribution were time-consuming and laborious.

[1335] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1336] In this invention, the server includes means for inputting information about a book that the user wants to read, means for receiving the book information input by the user and searching for the book in a data store, means for summarizing the entire text of the searched book, means for inputting the generated summary into a natural language processing engine, means for inputting the generated summary into a speech synthesis engine and converting it into audio data, means for temporarily storing the audio data and generating a download link thereof, and means for delivering the download link to the user's terminal. This makes it possible for the user to efficiently and quickly obtain an audio summary of the book.

[1337] A "user" is an individual or organization that uses the system to obtain book summaries and audio data.

[1338] "Book information" refers to information necessary to uniquely identify a book, such as the book's title, author's name, and ISBN code.

[1339] A "data store" is a storage device or database system used to store and manage book data.

[1340] "The full text of the book" refers to all the text and data contained in the book.

[1341] A "summary" refers to a concise compilation of the most important parts and key points extracted from the entire text of a book.

[1342] A "natural language processing engine" is software or algorithms that analyze text data and perform tasks such as summarizing, translating, and other natural language processing.

[1343] A "speech synthesis engine" is a software or hardware system that takes text data as input and converts it into speech data.

[1344] "Audio data" refers to audio files in an audible format, generated by a speech synthesis engine.

[1345] "Temporary storage" refers to the process of saving data only for a certain period, with deletion or updating performed as needed.

[1346] A "download link" refers to a URL or path used to obtain audio data or other data.

[1347] A "terminal" is an electronic device used by a user to access and operate a system.

[1348] This invention provides a system that enables users to efficiently understand the content of books even in their busy lives. This system has the function of automatically generating a book summary and providing it in audio format. Specific embodiments of this system are described in detail below.

[1349] System Configuration

[1350] This system primarily consists of a server and terminals, and users obtain book summaries as audio via the terminals. The system is comprised of the following hardware and software combinations.

[1351] hardware

[1352] The server includes a data store (e.g., a MySQL database), an NLP engine, and a speech synthesis engine.

[1353] Device: An electronic device used by a user, such as a smartphone, tablet, or personal computer.

[1354] software

[1355] Natural language processing engine (e.g., GPT-3)

[1356] Text-to-speech engine (e.g., Google Text-to-Speech)

[1357] Database management system (e.g., MySQL)

[1358] HTTP request processing libraries (e.g., Axios)

[1359] Operation flow

[1360] The operation flow of this system is as follows:

[1361] 1. Book Selection

[1362] The user enters the title or ISBN code of the book they want to read through the terminal's interface. For example, if the user wants to read "a famous novel," they enter the title in the terminal's search bar and press the "Search" button.

[1363] 2. Searching for book data

[1364] The terminal sends the book information entered by the user to the server. Based on the received information, the server searches its data store for matching book data. As a result, it sends back the full text of "a famous novel" and related metadata to the terminal.

[1365] 3. Automatic summary generation

[1366] The server passes the full text of the book to a natural language processing engine, which extracts the important parts and generates a short summary. For example, it passes a prompt sentence like the following to the natural language processing engine:

[1367] Please summarize the following text: "The complete text of a famous novel."

[1368] 4. Generating audio files

[1369] The server inputs the generated summary into the speech synthesis engine and converts it into speech data. Specifically, it uses prompts like the following:

[1370] "Please convert this summary into an audio file."

[1371] The generated audio file will be temporarily saved.

[1372] 5. Audio distribution

[1373] The server generates a download link for the audio file and notifies the terminal of that link. For example, it sends the user a message saying, "You can listen to an audio summary of a famous novel." The terminal then displays the received download link on its interface.

[1374] 6. Audio Playback

[1375] The user selects the audio playback option through the device's notifications or interface. When the user presses the "Play" button, the device retrieves the audio file from the download link and plays it using the built-in player.

[1376] 7. Add to favorites and receive notifications

[1377] Users can add books they want to read to their favorites list. When a new summary is generated, the server notifies the user via email or app notification.

[1378] This system enables users to efficiently understand the content of books even in their busy lives, and by automating each processing step, it achieves simple access and use. For example, it can be used in the same process when a user is trying to read the latest business book, "The Road to Success."

[1379] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1380] Step 1:

[1381] User: Enter information about the book you want to read. Specifically, the user enters the book title and ISBN code through the terminal interface and presses the "Search" button. This entered information is then sent to the next processing step.

[1382] Step 2:

[1383] Terminal: Sends the book information entered by the user to the server. Specifically, it sends the book information to the server using an HTTP request. This operation passes the input data to the server.

[1384] Step 3:

[1385] Server: Searches the data store based on the received book information. Specifically, it queries the data store (e.g., a MySQL database) based on the received book information and retrieves the full text and metadata of matching books. This query process outputs the full text and metadata.

[1386] Step 4:

[1387] Server: The server passes the full text obtained as search results to a natural language processing (NLP) engine to generate a summary. Specifically, it inputs the full text into the NLP engine using the prompt message, "Summarize the following text: The full text of 'Book Title'." The NLP engine performs text analysis, extracts the important parts, and generates a summary. This summary is then passed to the next step.

[1388] Step 5:

[1389] Server: Inputs the generated summary into a text-to-speech engine and converts it into audio data. Specifically, it inputs the summary into a text-to-speech engine (e.g., Google Text-to-Speech) using a prompt message such as "Please convert this summary into an audio file." The text-to-speech engine converts the text into speech and generates an audio file in MP3 format. This audio file is then passed on to the next step.

[1390] Step 6:

[1391] Server: Temporarily stores the audio file and generates a download link. Specifically, it uploads the generated audio file to cloud storage and generates a download link. This link is then sent to the next step.

[1392] Step 7:

[1393] Device: Notifies the user of the generated download link. Specifically, the device's interface and notification system display the link along with the message, "You can listen to an audio summary of the book title." When the user clicks this link, they proceed to the next step.

[1394] Step 8:

[1395] User: Selects the option to play the audio file. Specifically, the user clicks the received link and plays it using the device's built-in player. This playback action allows the user to listen to the summarized content in audio.

[1396] Step 9:

[1397] User: Adds books they want to read to their favorites list. Specifically, they press the "Add to Favorites" button on their device. This action saves the book's information to the user's personal data store.

[1398] Step 10:

[1399] Server: When a new summary is generated for a book added to the favorites list, a notification is sent. Specifically, when new summary data is generated, the user is notified via email or app notification. This notification allows the user to check the new information immediately.

[1400] (Application Example 1)

[1401] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1402] For users leading busy lives, there is a need for efficient ways to understand the content of books. However, traditional methods of providing book summaries and audiobooks sometimes make it difficult for users to quickly obtain the information they are looking for. Furthermore, there is a lack of systems that easily recommend new reading material and provide easy access to desired reading. A new system is needed to address these challenges.

[1403] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1404] In this invention, the server includes means for the user to input information about a reading material, means for receiving the information about the reading material input by the user and searching for the reading material in a database, means for summarizing the entire text of the searched reading material, means for converting the summarized content of the reading material into an audio file, means for distributing the audio file to the user's terminal device, and means for presenting recommended reading material based on the information about the reading material input by the user. As a result, the user can not only quickly listen to a summary in audio format based on the information they have input, but also receive recommendations for related new reading material, enabling them to understand the content of books more easily and efficiently.

[1405] A "user" is an individual or organization that uses the system to input information about reading material and receives the retrieved summaries and audio files.

[1406] "Reading materials" is a general term for information media provided in text format, such as books, magazines, articles, and ebooks.

[1407] "Means of input" refers to methods of providing information to a system using keyboards, voice input, touchscreens, etc.

[1408] "Means of receiving" refers to the function by which a server retrieves information sent by a user.

[1409] "Searching methods" refer to the function of finding relevant reading material from a database based on the received information.

[1410] "Full text" refers to text data that corresponds to the entire text of the searched reading material.

[1411] "Methods of summarization" refer to the function of extracting important parts from the entire text and summarizing them concisely.

[1412] "Means of converting to audio files" refers to a function that converts summarized text into audio data.

[1413] "Terminal devices" refer to electronic devices such as smartphones, tablets, and computers.

[1414] "Means of distribution" refers to the function of sending the generated audio file to the user's terminal device.

[1415] "Means for suggesting recommended reading material" refers to a function that suggests new reading material related to the user's input information.

[1416] The "playback option" refers to the choices a user can make to listen to an audio file.

[1417] "Adding to a favorites list" refers to a function that allows users to save reading material to a list so they can easily access it later.

[1418] "Means of sending notifications" refers to a function that allows the server to notify users when a new summary is generated.

[1419] This invention is a system designed to enable users to efficiently understand the content of reading materials. This system consists of a server and terminal devices, allowing users to input book information and obtain summarized content in audio format. A specific embodiment of this system is described below.

[1420] Hardware and software to be used

[1421] Hardware:

[1422] Terminal devices: Smartphones, tablets, computers, smart glasses, head-mounted displays, etc.

[1423] Server: A high-performance computer with internet connectivity.

[1424] software:

[1425] Python programming language

[1426] requests library: Used for communication between servers and terminal devices.

[1427] gTTS (Google Text-to-Speech) library: Converts text to speech.

[1428] playsound library: Plays the generated audio file.

[1429] System operation

[1430] Book Search

[1431] The user inputs the title and ISBN code of the book they want to read through the interface of the terminal device, and sends that information to the server. For example, the user inputs the title of "a famous novel."

[1432] Book summary generation

[1433] The server searches the database based on the input information and retrieves data on matching reading materials. The retrieved complete text is then summarized using a natural language processing engine. This summarization process extracts the main points and important parts of the text and condenses them into a concise form.

[1434] Audio file generation and distribution

[1435] The gTTS library is used to convert the generated summary text into an audio file. The generated audio file is delivered from the server to the terminal device, and the user can understand the summary by listening to it. For example, the content of a summarized "famous novel" is delivered as an audio file.

[1436] Recommended reading material

[1437] Based on the information entered about the reading material, the server generates and presents relevant recommended reading materials to the user. This allows the user to discover new and interesting reading material.

[1438] Specific example

[1439] For example, if a user enters the title "I Am a Cat," the system searches the database for the book, summarizes the entire text, and converts it into an audio file. This audio file is then delivered to the user's terminal device, allowing them to quickly listen to the content of "I Am a Cat." In addition, "Other Works by Soseki" are presented as recommended reading material.

[1440] Example of a prompt

[1441] Please enter the book title: I Am a Cat

[1442] Thus, the system of the present invention is a system that can efficiently provide summaries of reading materials and improve the reading experience by combining natural language processing technology and speech synthesis technology.

[1443] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1444] Step 1:

[1445] The user inputs information about the book they are reading from a terminal device. The input information, such as the book title and ISBN code, is transmitted to the server through the terminal device's interface.

[1446] Input: Book title and ISBN code

[1447] Output: Search information sent to the server

[1448] Step 2:

[1449] The server searches the database based on the information it receives about the reading material. The database stores the complete text of numerous reading materials and associated metadata.

[1450] Input: Search information (book title and ISBN code)

[1451] Output: Data on matching reading materials

[1452] Step 3:

[1453] The server passes the entire text retrieved as search results to a natural language processing (NLP) engine, which then generates a summary. The NLP engine extracts the important parts and creates a compact summary text.

[1454] Input: Entire text

[1455] Output: Summary text

[1456] Step 4:

[1457] The server inputs the generated summary text into a speech synthesis engine (e.g., the gTTS library) and converts it into an audio file. The audio file format commonly used is MP3.

[1458] Input: Summary text

[1459] Output: Audio file (MP3 format)

[1460] Step 5:

[1461] The server saves the generated audio file to temporary storage and generates a download link. The server then sends this link to the terminal device.

[1462] Input: Audio file (MP3 format)

[1463] Output: Download link

[1464] Step 6:

[1465] The user receives a download link from the terminal device, retrieves the audio file, and plays it on the terminal device's player. The terminal device then displays notifications and playback options.

[1466] Input: Download link

[1467] Output: Playback of audio file

[1468] Step 7:

[1469] Users add books they want to read to their favorites list. The server notifies the user via email or app notification when a new summary is generated.

[1470] Input: Registration information for your favorites list

[1471] Output: Notification of new summary

[1472] Step 8:

[1473] The server generates recommended reading material based on the user's input and presents it to the terminal device. Similar genres and authors are considered in the recommendations.

[1474] Input: User input information

[1475] Output: Presentation of recommended reading material

[1476] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1477] ---

[1478] This invention provides a system that enables users to efficiently understand the content of books even in their busy lives. This system combines a function that automatically generates book summaries and provides them in audio format with an emotion engine that recognizes the user's emotions.

[1479] System Overview

[1480] This system consists of a server, terminals, and an emotion engine. Users can input book information using the terminals and receive summarized content in audio. Furthermore, the emotion engine can suggest appropriate books and adjust summaries based on the user's emotions.

[1481] 1. Book Selection

[1482] User: Enter the title or ISBN code of the book you want to read using the terminal's interface. For example, if the user wants to read "a famous novel," they would enter the title in the terminal's search bar.

[1483] 2. Searching for book data

[1484] Terminal: Sends the book information entered by the user to the server.

[1485] Server: Based on the received information, it searches the database and finds data for matching books. It then sends the found data back to the terminal. For example, it might send the full text and related metadata of a famous novel to the terminal.

[1486] 3. Automatic summary generation

[1487] Server: The full text of a book is passed to a natural language processing (NLP) engine, which extracts the important parts and generates a summary. For example, the NLP engine extracts the key plot points and characters from the full text of a famous novel and summarizes them concisely.

[1488] 4. Adjusting emotion recognition and summarization

[1489] Server: Uses an emotion engine that analyzes the user's facial expressions, voice tone, input text, etc., to recognize the user's current emotional state. For example, the emotion engine might detect a happy emotion from the user's voice tone.

[1490] Server: Based on recognized emotions, adjusts the content and tone of the summary and audio file. For example, the emotion engine changes the tone of the summary according to the user's emotions, adjusting it to emphasize positive content.

[1491] 5. Generating audio files

[1492] Server: Inputs the adjusted summary into the speech synthesis engine and converts it into audio data. Specifically, the speech synthesis engine converts the summary into an MP3 audio file with an emotion-appropriate tone.

[1493] 6. Audio distribution

[1494] Server: Saves the generated audio files to temporary storage and generates a download link.

[1495] Device: Presents the received download link to the user. For example, a notification might appear stating, "You can listen to an audio summary of a famous novel."

[1496] 7. Audio Playback

[1497] User: Select an audio playback option from the device's notifications or interface. For example, the user presses the "Play" button.

[1498] Device: Download the audio file from the download link and play the audio using the built-in player.

[1499] 8. Add to favorites and receive notifications

[1500] User: Adds books they want to read to their favorites list. For example, the user presses the "Add to Favorites" button.

[1501] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[1502] In this way, the system of the present invention enables users to efficiently understand the content of books even in their busy lives, and further provides a personalized experience through the emotion engine. Each processing step is automated, making it easy for users to access and use. While specific examples have been provided, the system is not limited to these and can be applied in a variety of ways.

[1503] The following describes the processing flow.

[1504] ---

[1505] Step 1:

[1506] User: Open the dedicated application or webpage on your device and enter the title or ISBN code of the book you want to read. For example, if the user wants to read "a famous novel," they would enter the title in the search bar on their device.

[1507] Step 2:

[1508] Terminal: Sends the book information entered by the user to the server.

[1509] Step 3:

[1510] Server: Based on the title or ISBN code of the received book, the server searches the database for data on matching books.

[1511] Step 4:

[1512] Server: Returns the full text of the searched book and related metadata to the terminal.

[1513] Step 5:

[1514] Server: The full text of a book is passed to a natural language processing (NLP) engine, which extracts the important parts and generates a summary. For example, the NLP engine extracts the key plot points and characters from the full text of a famous novel and summarizes them concisely.

[1515] Step 6:

[1516] Server: After the summary is generated, an emotion engine analyzes the user's facial expressions, voice tone, input text, etc., to recognize the user's emotional state. For example, the emotion engine might detect a cheerful emotion from the user's voice tone.

[1517] Step 7:

[1518] Server: Adjusts the content and tone of the summary and audio file based on recognized emotions. For example, the emotion engine might change the tone of the summary to a positive one based on the user's positive emotions.

[1519] Step 8:

[1520] Server: Inputs the adjusted summary into the speech synthesis engine and converts it into audio data. Specifically, the speech synthesis engine converts the summary into an MP3 audio file with an emotion-appropriate tone.

[1521] Step 9:

[1522] Server: Saves the generated audio files to temporary storage and generates a download link.

[1523] Step 10:

[1524] Server: Sends a download link to the device.

[1525] Step 11:

[1526] Terminal: Receives the sent download link and displays a notification to the user stating, "The summary is available for audio playback."

[1527] Step 12:

[1528] User: Select the audio playback option in the device's notification or interface and press the "Play" button.

[1529] Step 13:

[1530] Device: Download the audio file from the download link and play the audio using the built-in player.

[1531] Step 14:

[1532] User: Add books you want to read to your favorites list as needed. For example, the user presses the "Add to Favorites" button.

[1533] Step 15:

[1534] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[1535] ---

[1536] The above is a detailed explanation of the processing steps. This system automates the entire process from book information input to audio playback, and further provides a personalized experience through its emotion engine. Users can easily access and use it.

[1537] (Example 2)

[1538] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1539] In modern society, users are required to consume information and content efficiently amidst their busy lives. However, traditional methods present challenges such as the lack of time to read long texts like books and the effort required to understand summarized information. Furthermore, providing personalized information tailored to the user's current emotional state is difficult.

[1540] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting information of content that the user wants to read, means for receiving the information of content input by the user and searching for the content from a database, means for automatically summarizing the full text of the searched content, means for adjusting the summarized content based on the user's emotional state, means for converting the adjusted content into an audio file, and means for delivering the audio file to the user's device. As a result, even in a busy life, the user can efficiently understand the main points of the content and receive personalized information tailored to their emotional state.

[1541] "User" refers to an individual or legal entity that intends to use this system to search for, summarize, or convert content into audio.

[1542] "Content" refers to information such as books, articles, and reports, and includes informational materials that users wish to read.

[1543] "Means of inputting information" refers to the interface that allows users to input information such as the content title, author name, and ISBN code into their device.

[1544] A "database" refers to a collection of information that stores, manages, and allows for quick searching of content metadata and full text.

[1545] "Methods for automatic summarization" refer to algorithms or engines that use natural language processing technology to extract key parts from the full text of content and generate a concise summary.

[1546] "Means of adjustment based on emotional state" refers to a function that analyzes the user's current emotions and adjusts the content and tone of information such as summaries and audio according to those emotions.

[1547] "Means of converting to audio files" refers to speech synthesis engines and software that convert automatically summarized text information into audio data.

[1548] "Device" refers to the electronic device that a user uses to access this system, and includes smartphones, tablets, personal computers, etc.

[1549] "Means of distribution" refers to the function of transmitting generated audio files over a network in order to provide them to users.

[1550] This invention is a system designed to enable users to efficiently understand content even in their busy lives. This system combines a function that automatically generates a summary of the content and provides it in audio format with an emotion engine that recognizes the user's emotions.

[1551] System Configuration

[1552] This system consists of a server, terminals, and an emotion engine. Users can input content information using the terminals and receive summarized content in audio format. Furthermore, the emotion engine can suggest appropriate content and adjust summaries based on the user's emotions.

[1553] Hardware and software to be used

[1554] 1. Servers: Use cloud servers or physical servers with high computing power.

[1555] Database: Use database solutions such as AWS RDS or MongoDB.

[1556] Natural language processing engine: Uses generative AI models such as OpenAI GPT-3.

[1557] Emotion recognition engine: Uses emotion recognition services such as the Microsoft Azure Emotion API.

[1558] Text-to-speech engine: Uses a text-to-speech service such as the Google Text-to-Speech API.

[1559] 2. Device: A smartphone, tablet, or personal computer used by the user.

[1560] User interface: Web browser or dedicated application.

[1561] Program processing flow

[1562] The server receives content information entered by the user, searches the database based on that information, and retrieves the full text of the relevant content. The retrieved full text is then passed to a natural language processing engine, which automatically generates a summary. Furthermore, based on the user's sentiment data obtained from the device, an emotion recognition engine is used to adjust the summary content and voice tone. The adjusted summary is input to a speech synthesis engine and generated as an MP3 audio file. This audio file is stored in temporary storage by the server, and a download link is generated. Finally, the link is sent to the user's device, and the user can play the audio file from this link.

[1563] Specific example

[1564] For example, if a user wants to read a famous novel, they enter the title into the search bar on their device. This information is sent to the server, which retrieves the full text of the novel from its database. The retrieved text is summarized by a generative AI model. Based on this summary, an emotion recognition engine recognizes the user's emotions from their facial expressions and voice tone, and adjusts the content and tone of the summary. The adjusted summary is converted into an MP3 audio file by a speech synthesis engine and provided to the user's device.

[1565] Example of a prompt

[1566] "Please provide an audio summary of Natsume Soseki's 'Kokoro,' with tone adjustments based on emotional recognition."

[1567] "Please automatically summarize the latest business books and generate audio files with tone adjustments based on emotion recognition."

[1568] In this way, the system of the present invention enables users to efficiently understand the key points of content even in their busy lives and to provide personalized information tailored to their emotional state.

[1569] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1570] Step 1: Enter content information

[1571] User: Enter the title or ISBN code of the content you want to read from the terminal's interface. For example, if the user wants to read "a famous novel," they would enter the title in the terminal's search bar.

[1572] Input: Title and ISBN code entered by the user.

[1573] Output: Input data is sent to the server.

[1574] Specific operation: The terminal interface displays "Please enter the title or ISBN of the content." The user enters the information using the keyboard and then clicks the "Search" button.

[1575] Step 2: Search for content data

[1576] Terminal: Sends information about the content entered by the user to the server.

[1577] Server: Based on the received information, it searches a database (e.g., AWS RDS or MongoDB) and finds data with matching content. It then sends the corresponding data back to the terminal in JSON format.

[1578] Input: Entered content information.

[1579] Output: The full text and associated metadata of the relevant content retrieved from the database.

[1580] Specific operation: The terminal sends input data to the server using a REST API. The server executes SQL or NoSQL queries against the database, retrieves the relevant data, converts it to JSON format, and sends it back to the terminal.

[1581] Step 3: Generate an automated summary

[1582] Server: The full text of the content is passed to a natural language processing (NLP) engine (e.g., OpenAI GPT-3 model), which extracts the important parts and generates a summary.

[1583] Input: Full text (JSON format).

[1584] Output: Summarized text.

[1585] Specific operation: The server inputs the full text of the retrieved content into the AI ​​model along with the prompt, "Please summarize:". The AI ​​model analyzes the main parts of the text, extracts the important parts, and generates a summary. The generated summary is saved as a text file.

[1586] Step 4: Adjusting Emotion Recognition and Summarization

[1587] User: While generating a summary, the device's built-in camera and microphone are used to collect user emotion data. For example, when the user is smiling while looking at the screen.

[1588] Server: Uses an emotion engine (e.g., Microsoft Azure Emotion API) that analyzes the user's facial expressions, voice tone, input text, etc., to recognize the user's current emotional state.

[1589] Input: User's facial expression, voice tone, and input text.

[1590] Output: Recognized emotion data.

[1591] Specific operation: The user's device captures facial expressions with its camera and collects audio with its microphone. This data is sent to a server, which uses an emotion engine to send a prompt message saying, "Please recognize the emotion." Based on the recognized emotion data, the content and tone of the summary are adjusted.

[1592] Step 5: Generate audio files

[1593] Server: Inputs the refined summary into a text-to-speech engine (e.g., Google Text-to-Speech API) and converts it into audio data.

[1594] Input: Adjusted summary text.

[1595] Output: MP3 audio file.

[1596] Specific operation: The server inputs the adjusted summary into the speech synthesis engine, along with the prompt, "Please convert this text into speech, including emotional tone." The generated audio data is saved to the server as an MP3 file.

[1597] Step 6: Audio Distribution

[1598] Server: Saves the generated audio files to temporary storage (e.g., Amazon S3) and generates a download link.

[1599] Terminal: Presents the received download link to the user.

[1600] Input: MP3 audio file.

[1601] Output: Download link.

[1602] Specific operation: The server uploads the audio file to temporary storage and generates a public download link. This link is sent to the device, and the device notifies the user of the link.

[1603] Step 7: Play back audio

[1604] User: Select an audio playback option from the device's notifications or interface. For example, the user presses the "Play" button.

[1605] Device: Download the audio file from the download link and play the audio using the built-in player.

[1606] Input: Download link.

[1607] Output: Audio played on the device.

[1608] Specific operation: When the user presses the "Play" button, the device retrieves the audio file from the download link, launches the media player, and plays the audio.

[1609] Step 8: Add to favorites and receive notifications

[1610] User: Adds content they want to read to their favorites list. For example, the user presses the "Add to Favorites" button.

[1611] Server: Configure settings to notify users via email or app notifications when new summaries are generated.

[1612] Input: Information about the content to add to your favorites list.

[1613] Output: Notification when a new summary is generated.

[1614] Specific operation: When a user presses the "Add to Favorites" button, information about that content is sent to the server and stored in the database. The server sends an email or app notification to the user when a new summary is added.

[1615] (Application Example 2)

[1616] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1617] In modern society, users lead busy lives and find it difficult to dedicate sufficient time to reading books. Therefore, there is a need for a system that allows for efficient comprehension of book content. However, existing systems merely provide summaries and lack personalization based on user emotions, resulting in low satisfaction. Furthermore, the audio summaries do not adapt to the user's state of mind, leading to insufficient information comprehension and engagement.

[1618] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting information of a book that the user wants to read, means for receiving the book information input by the user and searching for the book in a database, means for summarizing the full text of the searched book, means for converting the summarized book content into an audio file, means for distributing the audio file to the user's terminal, means for recognizing the user's emotions, and means for adjusting the summary content and tone based on the recognized emotions. As a result, users can efficiently understand the content of books even in their busy lives, and receive personalized book summaries based on their emotions in audio format, enabling a more satisfying information provision.

[1619] "User" refers to a person who uses the system to obtain summary information about books.

[1620] "Book information" refers to data used to identify a book, such as the book's title, ISBN code, and author's name.

[1621] A "database" refers to a collection of information that stores the full text of a book and related metadata.

[1622] "Full text" refers to the entire body of text in a book.

[1623] A "summary" refers to information that has been simplified by extracting the most important parts from the full text of a book.

[1624] An "audio file" refers to a digital file in which the content of a summarized book has been converted into audio data.

[1625] "Device" refers to electronic devices used by users, such as smartphones, smart glasses, and head-mounted displays.

[1626] "Emotion recognition" refers to technology that detects and analyzes the emotional state of a user.

[1627] "Tone" refers to the quality or tone of voice, and means a characteristic of voice that is altered based on emotion.

[1628] "Distribution" refers to the act of sending data from a server to a terminal.

[1629] "Personalization" refers to customization tailored to the user's specific needs and preferences.

[1630] Modes for carrying out the invention

[1631] System Overview

[1632] The system of this invention consists of a server, a terminal, and an emotion engine that recognizes the user's emotions. The user can use the terminal to input information about a specific book and obtain a summary of that book as an audio file. Furthermore, the emotion engine is used to adjust the summary according to the user's emotions.

[1633] Hardware and software used

[1634] Server: AWS (Amazon Web Services) or GCP (Google Cloud Platform)

[1635] Database: PostgreSQL or MySQL

[1636] NLP engine: SpaCy or BERT, both Python libraries.

[1637] Emotion Engine: An emotion recognition model using OpenCV and TensorFlow

[1638] Text-to-speech engine: Amazon Polly or Google Text-to-Speech

[1639] Devices: Smartphones, smart glasses, head-mounted displays

[1640] Explanation of the program's processing

[1641] 1. Book Selection

[1642] The user enters the title or ISBN code of the book they want to read through the terminal's interface. For example, if the user wants to read the book "1984," they would enter "1984" into the terminal's search bar.

[1643] 2. Searching for book data

[1644] The server receives the book information entered by the user, searches its database, and identifies the full text of any matching books. It then sends this result back to the terminal.

[1645] 3. Automatic summary generation

[1646] The server uses an NLP engine (e.g., spaCy or BERT) to extract key parts from the full text of the searched books and generate summaries.

[1647] 4. Emotion recognition

[1648] The device uses its camera and microphone to capture the user's facial expressions and voice tone, which are then analyzed by an emotion engine (e.g., OpenCV and TensorFlow models).

[1649] 5. Summarizing and adjusting the tone

[1650] Based on the analyzed sentiment data, the server adjusts the summary content and voice tone to personalize each user. This tone adjustment is performed during the voice file generation process.

[1651] 6. Generating audio files

[1652] The server uses a text-to-speech engine (e.g., Amazon Polly or Google Text-to-Speech) to convert the refined summary into audio data.

[1653] 7. Audio distribution

[1654] The server uploads the generated audio files to cloud storage and provides a download link to the user's device. The user can then retrieve the audio files via this link and make them available for use.

[1655] Specific example

[1656] The user opens the application and types "1984". The camera captures the user's face and determines that they are relaxed. The server uses this emotional information to generate a summary of the book "1984" in a friendly tone and provides it as an audio file. This entire process allows the user to efficiently understand the book's content even in a busy environment.

[1657] Example of a prompt

[1658] "Summarize the key information from all the texts from 1984 in 50 characters or less."

[1659] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1660] Step 1:

[1661] The user enters the title and ISBN code of the book they want to read through the terminal's interface. For example, if the user wants to read "1984," they enter "1984" into the terminal's search bar. This action results in data transfer, where the terminal sends the entered book information to the server.

[1662] Step 2:

[1663] The server queries the database for book information received from the terminal and searches for matching book data (full text and metadata). During this process, the server uses the title and ISBN code as keys to execute the search query. It retrieves the full text of the relevant books as a search result and sends it back to the terminal.

[1664] Step 3:

[1665] The server passes the full text of the book, retrieved from the database, to a natural language processing (NLP) engine. The NLP engine analyzes the text, extracts key parts, and generates a summary. For example, the NLP model might create a summary of the book's main plot and characters in 50 characters or less. This summary is temporarily stored.

[1666] Step 4:

[1667] The device uses its built-in camera and microphone to capture the user's facial expressions and voice tone. The collected data is sent to a server in real time. For example, the camera can capture the user's smile and send it as input data to an emotion recognition model.

[1668] Step 5:

[1669] The server uses an emotion recognition model (e.g., OpenCV and TensorFlow) to analyze the user's emotional state. It recognizes the user's emotions (e.g., joy, sadness, surprise, etc.) from the video and audio provided as input data. This analysis result is used to refine the summary data as an emotional state.

[1670] Step 6:

[1671] The server adjusts the content and tone of the summary generated by the NLP engine based on the user's emotional state. For example, if the user's emotion is "relaxed," the summary's tone will be adjusted to softer language to make it more approachable. As a result of this step, the adjusted summary data is generated.

[1672] Step 7:

[1673] The server inputs the adjusted summary data into a text-to-speech engine (e.g., Amazon Polly or Google Text-to-Speech) to generate an audio file. The text-to-speech engine generates audio data that matches the summary content and emotional tone, and converts it into an MP3 audio file. This audio file is stored on the server.

[1674] Step 8:

[1675] The server uploads the generated audio file to cloud storage and generates a download link. This link is resent to the device, and the user receives a notification. By clicking the link, the user can download and listen to the audio file.

[1676] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1677] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1678] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1679] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1680] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1681] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1682] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1683] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1684] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1685] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1686] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1687] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1688] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1689] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1690] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1691] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1692] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1693] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1694] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1695] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1696] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1697] The following is further disclosed regarding the embodiments described above.

[1698] ---

[1699] (Claim 1)

[1700] A means for users to input information about the books they want to read,

[1701] A means for receiving information about a book entered by the user and searching for the book in a database,

[1702] A means of summarizing the full text of the searched books,

[1703] A means of converting the summarized content of a book into an audio file,

[1704] A means for distributing the aforementioned audio file to the user's terminal,

[1705] A system that includes this.

[1706] (Claim 2)

[1707] The system according to claim 1, further comprising means for a user to select an option to play an audio file.

[1708] (Claim 3)

[1709] A method for users to add books they want to read to their favorites list,

[1710] When a new summary is generated, a means of sending a notification,

[1711] The system according to claim 1, further comprising:

[1712] ---

[1713] "Example 1"

[1714] (Claim 1)

[1715] A means for users to input information about the books they want to read,

[1716] A means for receiving information about a book entered by the user and searching for the book in a data store,

[1717] A means of summarizing the full text of the searched books,

[1718] A means for inputting the generated summary into a natural language processing engine,

[1719] A means of inputting the generated summary into a speech synthesis engine and converting it into speech data,

[1720] A means for temporarily storing the aforementioned audio data and generating a download link thereof,

[1721] A means for delivering the aforementioned download link to the user's device,

[1722] A system that includes this.

[1723] (Claim 2)

[1724] The system according to claim 1, further comprising means for a user to select an option to play audio data.

[1725] (Claim 3)

[1726] A method for users to add books they want to read to their favorites list,

[1727] When a new summary is generated, a means of sending a notification,

[1728] The system according to claim 1, further comprising:

[1729] "Application Example 1"

[1730] (Claim 1)

[1731] A means for users to input information about the books they are reading,

[1732] A means for receiving information on reading material entered by the user and searching for said reading material in a database,

[1733] A means of summarizing the entire text of the searched reading material,

[1734] A means of converting the content of summarized reading material into an audio file,

[1735] Means for distributing the aforementioned audio file to the user's terminal device,

[1736] A method for suggesting reading material based on information about reading material entered by the user,

[1737] A system that includes this.

[1738] (Claim 2)

[1739] The system according to claim 1, further comprising means for a user to select an option to play an audio file.

[1740] (Claim 3)

[1741] A method for users to register books they want to read to their favorites list,

[1742] When a new summary is generated, a means of sending a notification,

[1743] The system according to claim 1, further comprising:

[1744] "Example 2 of combining an emotion engine"

[1745] (Claim 1)

[1746] A means for users to input information about the content they want to read,

[1747] A means for receiving information on content entered by the user and searching for said content in a database,

[1748] A method for automatically summarizing the full text of searched content,

[1749] A means of adjusting summarized content based on the user's emotional state,

[1750] means for converting the adjusted content into an audio file,

[1751] A means for delivering the aforementioned audio file to the user's device,

[1752] A system that includes this.

[1753] (Claim 2)

[1754] The system according to claim 1, further comprising means for a user to select an option to play an audio file.

[1755] (Claim 3)

[1756] A method for users to add content they want to read to their favorites list,

[1757] When a new summary is generated, a means of sending a notification,

[1758] The system according to claim 1, further comprising:

[1759] "Application example 2 of combining emotional engines"

[1760] (Claim 1)

[1761] A means for users to input information about the books they want to read,

[1762] A means for receiving information about a book entered by the user and searching for the book in a database,

[1763] A means of summarizing the full text of the searched books,

[1764] A means of converting the summarized content of a book into an audio file,

[1765] A means for distributing the aforementioned audio file to the user's terminal,

[1766] Means for recognizing the user's emotions,

[1767] A means of adjusting the content and tone of the summary based on perceived emotions,

[1768] A system that includes this.

[1769] (Claim 2)

[1770] The system according to claim 1, further comprising means for a user to select an option to play an audio file.

[1771] (Claim 3)

[1772] A method for users to add books they want to read to their favorites list,

[1773] When a new summary is generated, a means of sending a notification,

[1774] The system according to claim 1, further comprising: [Explanation of Symbols]

[1775] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for users to input information about the books they want to read, A means for receiving information about a book entered by the user and searching for the book in a database, A means of summarizing the full text of the searched books, A means of converting the summarized content of a book into an audio file, A means for distributing the aforementioned audio file to the user's terminal, A system that includes this.

2. The system according to claim 1, further comprising means for a user to select an option to play an audio file.

3. A method for users to add books they want to read to their favorites list, When a new summary is generated, a means of sending a notification, The system according to claim 1, further comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A