system
The system addresses the challenge of varying reading abilities by analyzing and simplifying ebook content for audio delivery, ensuring effective learning experiences through user-tailored and feedback-driven improvements.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-17
- Publication Date
- 2026-04-30
AI Technical Summary
The decline in reading ability among young people hinders effective knowledge acquisition, particularly with complex books and specialized content, and existing systems fail to provide tailored teaching materials at appropriate levels for individual users, lacking efficient mechanisms for continuous improvement based on user feedback.
A system that analyzes ebook content using natural language processing to identify complexity, applies a simplification algorithm based on user comprehension level, generates simplified text as audio via speech synthesis, and collects user feedback for continuous improvement, providing an adaptable learning experience.
Enables efficient and easy-to-understand learning experiences for diverse users by tailoring content to individual needs and continuously improving based on user feedback, facilitating learning in various situations, including those where reading is difficult.
Smart Images

Figure 2026071560000001_ABST
Abstract
Description
Technical Field
[0004] , ,
[0005] , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern society, the decline in the reading ability of young people is an important issue in learning efficiency and knowledge acquisition. This issue causes problems such as difficulty in understanding particularly complex books and specialized content, and as a result, limits the opportunities for knowledge absorption. Also, due to the variation in reading ability, it is difficult to provide teaching materials at an appropriate level for individual users.
Means for Solving the Problems
[0005] This invention provides a means for analyzing the content of ebooks using natural language processing technology and identifying the complexity of the text. Furthermore, it applies an algorithm that simplifies the text based on the reading comprehension level set by the user, enabling the generation of content tailored to individual needs. The simplified text is also provided as audio data using speech synthesis technology, accommodating users who have difficulty reading printed text. In addition, the system will be continuously improved based on collected user feedback. This makes it possible to provide an efficient and easy-to-understand learning experience for a wide range of users.
[0006] "Natural language processing means" refers to a device or program that has the function of processing data using techniques that analyze language data and understand the structure and meaning of sentences.
[0007] "E-book data" refers to the content of a book stored in digital format, which is data that can be displayed and manipulated on a computer or electronic device.
[0008] "User reading comprehension level" is a standard that represents the difficulty level of text that individual users can understand, and is an indicator used to adjust the complexity of the content provided by the system.
[0009] An "algorithm for simplifying text" refers to a computational method or procedure that adjusts the difficulty of grammar and vocabulary while preserving the meaning of the original text, making the content easier to understand.
[0010] "Speech synthesis means" refers to a device or program that uses technology to convert text data into speech output, and has the role of generating synthesized speech.
[0011] A "user terminal" refers to an electronic device that a user can directly operate, and is a device that enables the display and playback of content.
[0012] "Means of collecting feedback" refers to technologies or devices that systematically collect evaluations and opinions from users, and includes functions that help evaluate and improve the system. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] To implement this invention, a server, user terminals, and network infrastructure are primarily required. The server receives ebook data and analyzes its content using natural language processing technology. Specifically, the server analyzes the data, determines the complexity of the text, and identifies key topics and keywords.
[0035] Next, the server applies a simplification algorithm according to the reading comprehension level set by the user. This algorithm transforms the text into a simpler, more understandable format. For example, it replaces technical terms with plainer language and summarizes redundant parts.
[0036] The simplified text is then passed to a speech synthesis system. Here, the server converts the text data into synthesized speech. The synthesized speech data is configured to be clear and natural, ensuring that users can listen to it without stress.
[0037] The generated audio data and simplified text are transmitted to the user's terminal via the network. The terminal has an interface to provide the received data to the user, allowing playback, stopping, and rewinding of the audio. Users can also listen to the audio while viewing the simplified text. In this way, it becomes possible to provide efficient learning opportunities even for users who have difficulty reading printed text.
[0038] Furthermore, feedback provided by users after using the system is sent to the server via the terminal. The server analyzes the collected feedback and uses it to improve the system. This allows the service to continuously improve the user experience.
[0039] As a concrete example, suppose a user wants to read a specialized book. The server analyzes the data of that book and simplifies the content to make it easier for the user to understand. Furthermore, by providing this simplified content as audio, learning becomes possible even in situations where reading printed text is difficult, such as during commutes or while traveling. In this way, the invention realizes the provision of an efficient and high-quality learning experience to a diverse range of users.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The server retrieves the book data specified by the user's terminal from an external database or cloud storage. The server temporarily stores the retrieved data in text format.
[0043] Step 2:
[0044] The server uses natural language processing technology to analyze the text data of books, identifying the structure and complexity of the writing. It also extracts key topics and keywords to grasp the overall content.
[0045] Step 3:
[0046] The server applies a text simplification algorithm based on the reading comprehension level set by the user. The algorithm adjusts the difficulty of vocabulary according to the settings and reconstructs sentences into a more concise form.
[0047] Step 4:
[0048] The server passes simplified text data to a speech synthesis engine to generate synthesized speech. The generated speech is natural and adjusted to be easy for the user to understand.
[0049] Step 5:
[0050] The server sends the generated audio data and simplified text data to the user's device. Encryption technology is used for data transfer for security reasons.
[0051] Step 6:
[0052] The device provides an interface for playing back received audio data. Users can use the device to control playback, pause, and rewind of the audio. It also supports visual learning by displaying simplified text.
[0053] Step 7:
[0054] The user enters feedback via a terminal after using the system. The terminal collects the user's feedback and sends it to the server.
[0055] Step 8:
[0056] The server analyzes the feedback it receives and uses it to improve the service. This will further enhance the user experience in future use.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] In our information-driven society, various document data are available as ebooks, but many users find it difficult to understand complex specialized and academic books. In particular, there is a problem in that users with different reading comprehension levels have difficulty obtaining information tailored to their individual needs. Furthermore, while providing audio information is important for users who have difficulty absorbing information visually, conventional technology has limitations in quality. In addition, mechanisms for efficiently utilizing user feedback to improve the system are not yet fully established.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] In this invention, the server includes means for analyzing document data using an information processing device and automatically extracting the complexity of the text and the main topics; means for applying calculation means to simplify the document based on the user's reading ability; and means for generating audio information from the simplified document using a speech generation device. This enables the provision of documents tailored to different reading levels and the generation of high-quality audio information, as well as continuous improvement of the system by utilizing user feedback.
[0062] An "information processing device" is a device used to analyze data and identify and transform its contents.
[0063] "Document data" refers to data containing text information stored in electronic format.
[0064] "Complexity" refers to the difficulty of understanding the sentences within a document, or to its syntactic and lexical complexity.
[0065] "Main topic" refers to the most important or central topic or theme within the document data.
[0066] "User reading comprehension ability" refers to the degree to which individual users are able to read and understand documents.
[0067] "Computational means" refers to a method of executing algorithms or programs for performing specific processes or transformations.
[0068] A "speech generation device" is a device or system for converting text information into speech.
[0069] "Auditory information" refers to information in the form of sound that can be recognized by hearing, generated by a speech generation device.
[0070] "Feedback" refers to opinions and evaluations provided by users after using a system, and is used to improve the system.
[0071] An "external data storage device" is an external data storage facility accessed via a network, used to retrieve necessary data.
[0072] To implement this invention, a server, user terminals, and a communication network infrastructure are required. First, the server receives document data, such as ebooks, via the network. This data is provided in a common format, making it accessible to users.
[0073] Next, the server utilizes a natural language processing library (e.g., NLTK or spaCy) to analyze the received document data. Here, it identifies the complexity and main topics of the document and organizes the information. This allows the server to understand the document's structure and content, laying the foundation for providing content tailored to the user.
[0074] Subsequently, the server uses a computational method (algorithm) to generate a simplified document based on the reading comprehension level set by the user. The algorithm replaces complex technical terms with simpler language and summarizes redundant content. Techniques used in this process include document summarization algorithms and paraphrasing techniques.
[0075] The simplified text is then converted into speech information via a speech generator. Here, a speech synthesis API (e.g., Google® Text-to-Speech API or Amazon Polly) is used to generate natural and clear synthesized speech. This speech is particularly effective as a visually independent means of information transmission.
[0076] The generated audio information and simplified text are transmitted to the user's terminal via the network. The user's terminal receives this information, displays it for easy user access, and provides audio playback, pause, and rewind functions. For example, a user can listen to the audio information using earphones while commuting and simultaneously review the simplified text.
[0077] As a concrete example, a user who wants to understand the textbook "Digital Signal Processing" requests the data from the server. The server converts the "complex mathematical explanations" into "basic data analysis methods" and provides them to the user as audio. An example of a prompt used in this process would be, "Please convert this text into simple language that even a middle school student can understand."
[0078] In this way, the server and user terminals work together to create a system that provides efficient and easy-to-understand information.
[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0080] Step 1:
[0081] The server receives ebook data via the network based on user requests. It takes ebook file format data (e.g., EPUB, PDF) as input and outputs parseable text data. Specifically, the server downloads data from specified resources using network protocols.
[0082] Step 2:
[0083] The server analyzes received text data using natural language processing libraries (e.g., NLTK, spaCy). It takes text data as input and outputs metadata identifying sentence complexity and main topics. The server performs data processing such as tokenization, part-of-speech tagging, and noun phrase extraction. Specifically, the server executes a parsing algorithm to extract important keywords and concepts from the text.
[0084] Step 3:
[0085] The server simplifies text, taking into account the user's reading level. It takes parsed metadata and user reading level information as input and outputs simplified text. The server applies algorithms to convert complex terminology into simpler expressions and summarize redundant sentences. Specifically, the server uses parsing to reconstruct sentences into a more concise form.
[0086] Step 4:
[0087] The server converts simplified text into speech information using a speech generator. Simplified text is the input, and synthesized speech data is the output. The server uses a speech synthesis API (e.g., Google Text-to-Speech API) to convert the text into an audio file (e.g., MP3 format). Specifically, the server adjusts phonemes and intonation to generate natural and clear speech.
[0088] Step 5:
[0089] The server sends the generated audio information and simplified text to the user terminal over the network. The input consists of an audio file and simplified text, and the output is sent in a format that can be displayed and played on the user terminal. The server composes the data package and prepares it for transmission according to the transmission protocol. Specifically, the server binds to the destination address and transmits the data through the specified port.
[0090] Step 6:
[0091] The user terminal provides received audio and text information, enabling the user to listen to and view it. It receives data from the server as input and outputs an interactive display for the user. The terminal launches an audio player and text viewer, providing playback, pause, and rewind functions. Specifically, the terminal displays play and pause buttons on the user interface and outputs audio.
[0092] Step 7:
[0093] Users provide feedback after using the system. The input consists of the user's subjective evaluation and comments, and the output is feedback data sent to the server. Users submit their opinions by entering their evaluation through a feedback form on their terminal and pressing the submit button. Specifically, users can select evaluation items and leave comments in the text field.
[0094] Step 8:
[0095] The server analyzes the collected feedback to help improve the system. It receives feedback data from users as input and outputs information for developing system improvement strategies. The server uses text analysis techniques to extract common problems and improvement requests from the feedback. Specifically, the server stores the feedback data in a database and uses machine learning algorithms to analyze patterns.
[0096] (Application Example 1)
[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0098] In recent years, with the spread of e-books and digital information, there has been a growing demand for information to be acquired efficiently and in an easily understandable format by users at various levels. However, much of this information is specialized and complex, making it difficult for some users to comprehend. Furthermore, there is a need to efficiently acquire information even in situations where directly reading text is difficult, such as during commutes or while traveling. These problems are particularly pronounced when using specialized books or materials containing advanced content. Therefore, there is a need for systems that simplify e-books and digital information according to the user's level of understanding and provide it in audio format.
[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0100] In this invention, the server includes means for analyzing digital information data using natural language processing technology to automatically identify the difficulty level and important themes of the text, means for applying computational means to simplify the text according to the user's level of understanding, and means for creating audio information from the simplified text using speech generation technology. This enables personalized learning assistance for the user by simplifying the text and converting it to speech based on the digital information content selected by the user.
[0101] "Natural language processing technology" refers to the technology that enables computers to understand, interpret, and generate natural language, and is used to analyze the meaning of text.
[0102] "Digital information data" refers to a collection of information stored in digital format, including ebooks and online articles.
[0103] "Difficulty level of a text" is an indicator of the level of comprehension required for a text, and is evaluated based on factors such as the frequency of use of complex grammar and technical terms.
[0104] An "important theme" refers to a particularly important or central topic within a text or piece of information, requiring a deep understanding of its content.
[0105] "Comprehension level" is an indicator of how accurately a user can understand information, and it depends on the individual's knowledge level and experience.
[0106] "Computational means" include mathematical methods and algorithms for simplifying complex information according to the characteristics of the user.
[0107] "Speech generation technology" is a technology that converts text data into speech data and is used to achieve natural and clear speech output.
[0108] "Personalized learning support" is a method of providing learning support tailored to the user's needs and level, thereby offering an efficient learning experience.
[0109] In order to implement this invention, the combination of server, user terminal, and network infrastructure is crucial. The specific configuration is described below.
[0110] First, the server analyzes digital information data using natural language processing technology. The software used here includes natural language processing libraries such as SpaCy and Transformers. This allows for the automatic identification of the difficulty level and important themes of the text, enabling information processing tailored to specific users.
[0111] The analyzed data is simplified based on the user's level of understanding. This process applies computational methods tailored to the user's characteristics, transforming complex information into a simpler form. The simplified text is then converted into audio using speech generation technology such as Google Text-to-Speech, and presented to the user.
[0112] The converted audio information and simplified text are transmitted to the user terminal via the network. The user terminal has an interface for displaying the received information and playing the audio. This allows the user to easily obtain information both visually and aurally.
[0113] For example, if a user selects a specialized academic book, the server analyzes its content, simplifies it according to the user's level of understanding, and then provides it as audio that can be listened to even while commuting. In this way, users can effectively learn while on the go.
[0114] An example of a prompt message is shown below.
[0115] Text: 'Please enter a portion of a complex book here.'
[0116] User level: 'Intermediate'
[0117] Output: 'Simplified text'
[0118] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0119] Step 1:
[0120] The user selects and inputs digital information data into the terminal, and the terminal sends that data to the server. The input here is text data from ebooks or online articles, and the output is raw text data sent to the server. The terminal uses the HTTP protocol to send the data to the server.
[0121] Step 2:
[0122] The server parses the received text data. The input is raw text data, and the output is information identifying the difficulty level and important themes of the text. The server uses natural language processing libraries (e.g., SpaCy or Transformers) to identify the structure and main topics of the text and generate metadata for the data.
[0123] Step 3:
[0124] The server simplifies text data based on identification information and user comprehension settings. The input is identification information and user comprehension settings, and the output is simplified text. The server uses specific algorithms (e.g., generative AI models) to convert complex technical terms into simpler expressions and summarize redundant parts.
[0125] Step 4:
[0126] The server converts simplified text into audio data. The input is simplified text, and the output is audio data. The server uses speech generation technology (e.g., Google Text-to-Speech) to generate the text data as an audio file.
[0127] Step 5:
[0128] The server sends the generated audio data and simplified text to the terminal. The input is the audio data and simplified text, and the output is the terminal that receives this data. Since the server streams the data over the network, the user can obtain information in real time.
[0129] Step 6:
[0130] The user plays audio and views simplified text on their device. Input consists of audio data and text sent from the server, while output is information the user hears and sees. The device provides an audio player and text viewer to support the user's learning experience.
[0131] Step 7:
[0132] Users send feedback from their devices to the server. The input is user feedback information, and the output is data stored on the server. The device provides a feedback form and reports the user's experience to the server.
[0133] Step 8:
[0134] The server analyzes the collected feedback and uses it to improve the system. The input is feedback data provided by users, and the output is an optimized system configuration. By analyzing this feedback, the server periodically improves its generation algorithm, aiming to provide more effective information.
[0135] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0136] This invention is a system that provides an interactive learning environment tailored to the user's emotional state by combining an emotion engine. The server retrieves book data selected by the user from an external database and analyzes its contents using natural language processing technology.
[0137] After analysis, the server determines the complexity and identifies key topics, then simplifies the text based on the user's reading comprehension level. At this stage, the emotion engine activates, analyzing the user's facial expressions and voice through the user's device's camera and sensors to identify emotions.
[0138] Once an emotional state is identified, the server adjusts the parameters of its simplification algorithm to change the difficulty and expression of the text according to that emotion. For example, if the server determines that the user is tired, it will provide simpler expressions and shorter sentences. In addition, the tone and speed of the synthesized speech will change according to the user's emotion. If the user is excited, a faster tone will be used; if the user is relaxed, a slower tone will be used.
[0139] The generated audio data and refined text data are sent to the user's terminal and provided in a displayable and playable format. Users can learn efficiently by reading the simplified text or listening to the audio. Furthermore, when users input feedback into their terminal after using the system, the server utilizes this feedback to improve the algorithm and the accuracy of sentiment recognition.
[0140] For example, if the emotion engine detects stress while a user is reading a scientific textbook, the server will avoid technical jargon and simplify the grammar to make the content easier to understand. Speech synthesis will also be provided in a calmer tone, allowing users to absorb knowledge while reducing stress. This enables the invention to provide a flexible and personalized learning experience.
[0141] The following describes the processing flow.
[0142] Step 1:
[0143] The user selects the book they want to read through their device. This selection information is then sent to the server.
[0144] Step 2:
[0145] The server retrieves data for the selected books from an external database via an API and temporarily stores it in text format.
[0146] Step 3:
[0147] The server analyzes the book data using natural language processing techniques. The server identifies the complexity of the text and extracts key topics and keywords.
[0148] Step 4:
[0149] The device's built-in camera and sensors capture the user's facial expressions and voice data. The device then sends this data to an emotion engine.
[0150] Step 5:
[0151] The emotion engine analyzes the user's facial expressions and voice data to identify their current emotional state. It then sends this emotional state information to the server.
[0152] Step 6:
[0153] The server applies a text simplification algorithm based on the user's reading comprehension level and emotional state. The server then generates the simplified text.
[0154] Step 7:
[0155] The server inputs simplified text into a speech synthesis engine, which then synthesizes speech data with a tone and speed that matches the user's emotions.
[0156] Step 8:
[0157] The server sends the generated audio data and simplified text to the user's device.
[0158] Step 9:
[0159] The device provides the user with an interface for playing back received audio data. The user can play the audio and visually review the simplified text.
[0160] Step 10:
[0161] The user enters feedback into the terminal after using the system. The terminal sends the collected feedback information to the server.
[0162] Step 11:
[0163] The server analyzes feedback data and uses it to improve the entire system. It enhances the accuracy of algorithms and the sentiment engine to provide a better user experience.
[0164] (Example 2)
[0165] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0166] Traditional learning systems have faced challenges in providing learning materials that adequately consider the user's emotional state, making efficient learning difficult. Furthermore, the fixed difficulty level of documents prevents flexible adaptation to the user's reading comprehension level and understanding. Additionally, there was a lack of means to continuously improve the system based on the user's learning experience.
[0167] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0168] In this invention, the server includes means for analyzing document data using natural language processing means and simplifying the text based on the user's reading comprehension level; means for analyzing the user's emotional state in real time and dynamically adjusting the content based on that analysis; and means for collecting user feedback and using it to improve the system's accuracy. This enables a flexible and personalized learning experience.
[0169] "Natural language processing methods" are technologies that analyze document data to identify the complexity and main topics of the text.
[0170] "User reading comprehension level" is an indicator that shows how well a user can understand the content of a document.
[0171] "Processing to simplify text" refers to the process of adjusting the difficulty level of a document to suit the user.
[0172] "Emotional data" refers to information such as facial expressions and voice that is acquired in order to analyze the user's emotional state.
[0173] "Dynamic adjustment methods" refer to technologies that have the ability to change the presentation of content according to the user's emotions and state.
[0174] "Speech synthesis means" refers to technology that converts text data into speech data.
[0175] A "user terminal" is an input and output interface device used by a user.
[0176] "Feedback" refers to information about the system, including opinions and evaluations collected from users.
[0177] In order to implement this invention, the server, terminal, and user interfaces must work together in coordination.
[0178] The server first receives book requests from user terminals. Book data is retrieved from an external database and analyzed using natural language processing techniques. Generative AI models such as BERT and GPT are used for this analysis to identify the complexity and main topics of the text. Based on this, the server simplifies the text to suit the user's reading level.
[0179] Meanwhile, the user's device collects facial expressions and voice data in real time through its camera and microphone. This data is sent to a server and used for emotion recognition. This process utilizes common facial recognition APIs and voice analysis APIs.
[0180] Based on emotional data, the server dynamically adjusts text and voice expressions to match the user's emotional state. For example, if the user is stressed, technical jargon is eliminated and replaced with more concise and easily understandable language. Similarly, synthesized speech is generated with a tone and speed that matches the user's emotions.
[0181] The generated audio data and adjusted text data are sent to the user's terminal and provided visually and audibly. Through this, the user can obtain an optimized learning experience. Furthermore, as the user provides feedback during the learning process, the server uses it to improve the entire system.
[0182] For example, when a user is reading a physics textbook, the server simplifies the text and adjusts the audio data based on acquired book data and real-time sentiment analysis, allowing the user to absorb the knowledge in a relaxed state.
[0183] An example of a prompt might be, "Analyze the user's facial expressions and voice to estimate their emotions and provide appropriate learning materials."
[0184] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0185] Step 1:
[0186] The server receives a book selection request sent from the user's terminal. The input is the title or identifier of the book selected by the user. The server sends a query to an external database to retrieve the book data. The retrieved book data is then output.
[0187] Step 2:
[0188] The server analyzes acquired book data using natural language processing techniques. The input is the text data of the books. Generative AI models (e.g., BERT, GPT) are used to analyze the complexity of the text and identify key topics. The output provides text feature information based on the analysis results.
[0189] Step 3:
[0190] The device captures the user's facial expressions and voice data. Inputs include real-time video and audio data. The device sends this data to the server. The server receives raw sensor data based on facial expressions and voice as output.
[0191] Step 4:
[0192] The server analyzes emotional data transmitted from the terminal. Facial and voice data are used as input. Standard facial recognition and voice analysis technologies are applied to identify the user's emotional state. Information regarding the user's emotional state is generated as output.
[0193] Step 5:
[0194] The server dynamically adjusts the text of book data based on the analyzed emotional state. Inputs are user emotional information and book feature information. A simplification algorithm is applied, eliminating technical jargon as needed and converting the text into concise language. The output is the adjusted text data.
[0195] Step 6:
[0196] The server performs speech synthesis based on simplified text. The input consists of new text data and parameters to reflect the user's emotions. This results in output speech data generated with a tone and speed appropriate to the emotional state.
[0197] Step 7:
[0198] The server sends the adjusted text data and generated audio data to the user's terminal. The input is this data, which can then be displayed and played back visually and audibly on the terminal. The output is content in a format usable by the user.
[0199] Step 8:
[0200] After completing the learning process, users input feedback into their device. This feedback includes their opinions and evaluations. The device sends this feedback to a server, and the output is data that is used to improve the system.
[0201] (Application Example 2)
[0202] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0203] Traditional learning systems have a problem in that they cannot flexibly respond to the individual emotional states of users, making it difficult to provide an appropriate learning environment. In particular, there was no mechanism to adjust the content and expression to match the user's emotions when they were feeling stressed or fatigued. As a result, users may experience decreased learning efficiency and a decline in motivation to learn.
[0204] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0205] In this invention, the server includes means for analyzing e-book data using natural language processing means to automatically identify the complexity and main topics of the text, means for analyzing the user's facial expressions and vocal state using emotion recognition means to identify emotions, and means for adjusting the difficulty level and expression of the text according to the emotional state. This makes it possible to adjust the learning content and expression to suit the user's emotional state.
[0206] "Natural language processing methods" are technologies that analyze e-book data and automatically identify the complexity and main topics of the text.
[0207] "User comprehension level" is an indicator that shows how well a user can understand a text, and it is used to adjust the learning content.
[0208] "Speech formation means" refers to a technology that generates speech information from simplified text.
[0209] A "user device" is a terminal used by a user, and is a device that enables the display and playback of audio information and simplified text.
[0210] "Emotion recognition means" refers to technology that identifies emotions by analyzing the user's facial expressions and vocal state.
[0211] "Emotional state" refers to the temporary emotional state a user is experiencing, and it is information that influences the difficulty level and expression adjustments of the learning content.
[0212] This invention realizes a system that provides an interactive learning environment that responds to the user's emotional state. First, the server retrieves e-book data from an external data store. This data is automatically analyzed using natural language processing to identify its complexity and main topics. Based on this, an algorithm is applied to simplify the text based on the user's level of understanding.
[0213] Next, the device collects data on the user's facial expressions and voice, and identifies emotions using emotion recognition tools. This process utilizes APIs such as Google Cloud Vision API and voice analysis API to obtain real-time emotion information from the user. Based on this obtained emotion information, the server adjusts the difficulty and expression of the simplified text, and generates voice information from the adjusted text using speech generation tools.
[0214] As a concrete example, if a user is reading a physics article on their smartphone and the device detects signs of fatigue, the server will respond by converting the text into simpler language, avoiding complex terminology, and generating voice information in a calmer tone. In this case, a speech synthesis API such as Amazon Polly would be used.
[0215] Finally, this adjusted text and generated audio information are sent to the user's device and played back as both text and audio. Through this, the user can receive a personalized learning experience, improving learning efficiency and motivation. The system is designed so that the generating AI model makes optimal adjustments based on a prompt message such as, "How should the content be simplified and the voice tone adjusted if the user is judged to be tired?"
[0216] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0217] Step 1:
[0218] The server retrieves e-book data from an external data store. This data is passed to a natural language processing system to analyze the complexity of the text and its main topics. The information obtained from this analysis (degree of complexity and main topics) becomes the output.
[0219] Step 2:
[0220] Based on the output of Step 1, the server applies an algorithm to simplify the text to match the user's level of understanding. The input consists of complexity information and the user's level of understanding, and the server generates a simplified text based on this. The resulting simplified document becomes the output.
[0221] Step 3:
[0222] The device collects the user's facial expressions and voice through its camera and microphone, and transmits this data to an emotion recognition system for analysis. The input is the user's facial expressions and voice data, and the emotion recognition result is output.
[0223] Step 4:
[0224] The server uses the sentiment recognition results obtained in step 3 to adjust the difficulty level and expression of the simplified text. The input is the simplified text and the sentiment recognition results, and the output is the adjusted text.
[0225] Step 5:
[0226] The server converts the adjusted text into speech data using speech synthesis means. The input is the adjusted text, and the speech data is output by speech synthesis.
[0227] Step 6:
[0228] The terminal receives audio data and edited text sent from the server, and displays and plays them for the user. The input is the data received from the server, and the output is the presentation of visual and auditory information to the user.
[0229] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0230] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0231] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0232] [Second Embodiment]
[0233] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0234] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0235] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0236] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0237] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0238] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0239] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0240] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0241] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0242] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0243] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0244] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0245] To implement this invention, a server, user terminals, and network infrastructure are primarily required. The server receives ebook data and analyzes its content using natural language processing technology. Specifically, the server analyzes the data, determines the complexity of the text, and identifies key topics and keywords.
[0246] Next, the server applies a simplification algorithm according to the reading comprehension level set by the user. This algorithm transforms the text into a simpler, more understandable format. For example, it replaces technical terms with plainer language and summarizes redundant parts.
[0247] The simplified text is then passed to a speech synthesis system. Here, the server converts the text data into synthesized speech. The synthesized speech data is configured to be clear and natural, ensuring that users can listen to it without stress.
[0248] The generated audio data and simplified text are transmitted to the user's terminal via the network. The terminal has an interface to provide the received data to the user, allowing playback, stopping, and rewinding of the audio. Users can also listen to the audio while viewing the simplified text. In this way, it becomes possible to provide efficient learning opportunities even for users who have difficulty reading printed text.
[0249] Furthermore, feedback provided by users after using the system is sent to the server via the terminal. The server analyzes the collected feedback and uses it to improve the system. This allows the service to continuously improve the user experience.
[0250] As a concrete example, suppose a user wants to read a specialized book. The server analyzes the data of that book and simplifies the content to make it easier for the user to understand. Furthermore, by providing this simplified content as audio, learning becomes possible even in situations where reading printed text is difficult, such as during commutes or while traveling. In this way, the invention realizes the provision of an efficient and high-quality learning experience to a diverse range of users.
[0251] The following describes the processing flow.
[0252] Step 1:
[0253] The server retrieves the book data specified by the user's terminal from an external database or cloud storage. The server temporarily stores the retrieved data in text format.
[0254] Step 2:
[0255] The server uses natural language processing technology to analyze the text data of books, identifying the structure and complexity of the writing. It also extracts key topics and keywords to grasp the overall content.
[0256] Step 3:
[0257] The server applies a text simplification algorithm based on the reading comprehension level set by the user. The algorithm adjusts the difficulty of vocabulary according to the settings and reconstructs sentences into a more concise form.
[0258] Step 4:
[0259] The server passes simplified text data to a speech synthesis engine to generate synthesized speech. The generated speech is natural and adjusted to be easy for the user to understand.
[0260] Step 5:
[0261] The server sends the generated audio data and simplified text data to the user's device. Encryption technology is used for data transfer for security reasons.
[0262] Step 6:
[0263] The device provides an interface for playing back received audio data. Users can use the device to control playback, pause, and rewind of the audio. It also supports visual learning by displaying simplified text.
[0264] Step 7:
[0265] The user enters feedback via a terminal after using the system. The terminal collects the user's feedback and sends it to the server.
[0266] Step 8:
[0267] The server analyzes the feedback it receives and uses it to improve the service. This will further enhance the user experience in future use.
[0268] (Example 1)
[0269] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0270] In our information-driven society, various document data are available as ebooks, but many users find it difficult to understand complex specialized and academic books. In particular, there is a problem in that users with different reading comprehension levels have difficulty obtaining information tailored to their individual needs. Furthermore, while providing audio information is important for users who have difficulty absorbing information visually, conventional technology has limitations in quality. In addition, mechanisms for efficiently utilizing user feedback to improve the system are not yet fully established.
[0271] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0272] In this invention, the server includes means for analyzing document data using an information processing device and automatically extracting the complexity of the text and the main topics; means for applying calculation means to simplify the document based on the user's reading ability; and means for generating audio information from the simplified document using a speech generation device. This enables the provision of documents tailored to different reading levels and the generation of high-quality audio information, as well as continuous improvement of the system by utilizing user feedback.
[0273] An "information processing device" is a device used to analyze data and identify and transform its contents.
[0274] "Document data" refers to data containing text information stored in electronic format.
[0275] "Complexity" refers to the difficulty of understanding the sentences within a document, or to its syntactic and lexical complexity.
[0276] "Main topic" refers to the most important or central topic or theme within the document data.
[0277] "User reading comprehension ability" refers to the degree to which individual users are able to read and understand documents.
[0278] "Computational means" refers to a method of executing algorithms or programs for performing specific processes or transformations.
[0279] A "speech generation device" is a device or system for converting text information into speech.
[0280] "Auditory information" refers to information in the form of sound that can be recognized by hearing, generated by a speech generation device.
[0281] "Feedback" refers to opinions and evaluations provided by users after using a system, and is used to improve the system.
[0282] An "external data storage device" is an external data storage facility accessed via a network, used to retrieve necessary data.
[0283] To implement this invention, a server, user terminals, and a communication network infrastructure are required. First, the server receives document data, such as ebooks, via the network. This data is provided in a common format, making it accessible to users.
[0284] Next, the server utilizes a natural language processing library (e.g., NLTK or spaCy) to analyze the received document data. Here, it identifies the complexity of the document and the main topics, and organizes the information. As a result, the server understands the structure and content of the document, laying the foundation for providing content suitable for the user.
[0285] After that, based on the reading level set by the user, the server uses computational means (algorithms) to generate a simplified document. The algorithms replace difficult technical terms with more accessible words and summarize redundant content. Technologies used in this process include document summarization algorithms and paraphrasing techniques.
[0286] The simplified text is then converted into audio information via an audio generation device. Here, an audio synthesis API (e.g., Google Text-to-Speech API or Amazon Polly) is used to generate natural and clear synthetic speech. This audio is particularly effective as a means of information transmission that does not rely on vision.
[0287] The generated audio information and the simplified text are transmitted to the user terminal via the network. The user terminal receives this information, displays it for easy access by the user, and provides functions for playing, pausing, and rewinding the audio. For example, the user can listen to the audio information using earphones during commuting and simultaneously view the simplified text.
[0288] As a specific example, a user who wants to read the professional book "Digital Signal Processing" requests the data from the server. The server converts the "complex mathematical explanations" part into "basic data analysis methods" and provides it to the user as audio. An example of the prompt text used in this case is "Please convert this text into simple words so that even middle school students can understand it."
[0289] In this way, the server and the user terminal cooperate to build a system that realizes efficient and easy-to-understand information provision.
[0290] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0291] Step 1:
[0292] The server receives ebook data via the network based on user requests. It takes ebook file format data (e.g., EPUB, PDF) as input and outputs parseable text data. Specifically, the server downloads data from specified resources using network protocols.
[0293] Step 2:
[0294] The server analyzes received text data using natural language processing libraries (e.g., NLTK, spaCy). It takes text data as input and outputs metadata identifying sentence complexity and main topics. The server performs data processing such as tokenization, part-of-speech tagging, and noun phrase extraction. Specifically, the server executes a parsing algorithm to extract important keywords and concepts from the text.
[0295] Step 3:
[0296] The server simplifies text, taking into account the user's reading level. It takes parsed metadata and user reading level information as input and outputs simplified text. The server applies algorithms to convert complex terminology into simpler expressions and summarize redundant sentences. Specifically, the server uses parsing to reconstruct sentences into a more concise form.
[0297] Step 4:
[0298] The server converts simplified text into speech information using a speech generator. Simplified text is the input, and synthesized speech data is the output. The server uses a speech synthesis API (e.g., Google Text-to-Speech API) to convert the text into an audio file (e.g., MP3 format). Specifically, the server adjusts phonemes and intonation to generate natural and clear speech.
[0299] Step 5:
[0300] The server sends the generated audio information and simplified text to the user terminal over the network. The input consists of an audio file and simplified text, and the output is sent in a format that can be displayed and played on the user terminal. The server composes the data package and prepares it for transmission according to the transmission protocol. Specifically, the server binds to the destination address and transmits the data through the specified port.
[0301] Step 6:
[0302] The user terminal provides received audio and text information, enabling the user to listen to and view it. It receives data from the server as input and outputs an interactive display for the user. The terminal launches an audio player and text viewer, providing playback, pause, and rewind functions. Specifically, the terminal displays play and pause buttons on the user interface and outputs audio.
[0303] Step 7:
[0304] Users provide feedback after using the system. The input consists of the user's subjective evaluation and comments, and the output is feedback data sent to the server. Users submit their opinions by entering their evaluation through a feedback form on their terminal and pressing the submit button. Specifically, users can select evaluation items and leave comments in the text field.
[0305] Step 8:
[0306] The server analyzes the collected feedback and uses it to improve the system. There is feedback data sent from the user as input, and information for formulating strategies regarding system improvement is obtained as output. The server uses text analysis technology to extract common problems and improvement requirements from the opinions. As specific operations, the server stores the feedback data in a database and analyzes patterns using machine learning algorithms.
[0307] (Application Example 1)
[0308] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0309] In recent years, with the popularization of e-books and digital information, users at various levels are required to obtain information in an efficient and easy-to-understand manner. However, since many pieces of information have specialized and complex content, it is difficult for some users to understand. Also, there is a need to efficiently obtain information even in situations where it is difficult to directly read text information, such as during commuting or while moving. Such problems are particularly prominent when using materials containing specialized books or advanced content. Therefore, there is a demand for a system that simplifies e-books and digital information according to the user's level of understanding and provides it in audio.
[0310] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.
[0311] In this invention, the server includes means for analyzing digital information data using natural language processing technology to automatically identify the difficulty level and important themes of the text, means for applying computational means to simplify the text according to the user's level of understanding, and means for creating audio information from the simplified text using speech generation technology. This enables personalized learning assistance for the user by simplifying the text and converting it to speech based on the digital information content selected by the user.
[0312] "Natural language processing technology" refers to the technology that enables computers to understand, interpret, and generate natural language, and is used to analyze the meaning of text.
[0313] "Digital information data" refers to a collection of information stored in digital format, including ebooks and online articles.
[0314] "Difficulty level of a text" is an indicator of the level of comprehension required for a text, and is evaluated based on factors such as the frequency of use of complex grammar and technical terms.
[0315] An "important theme" refers to a particularly important or central topic within a text or piece of information, requiring a deep understanding of its content.
[0316] "Comprehension level" is an indicator of how accurately a user can understand information, and it depends on the individual's knowledge level and experience.
[0317] "Computational means" include mathematical methods and algorithms for simplifying complex information according to the characteristics of the user.
[0318] "Speech generation technology" is a technology that converts text data into speech data and is used to achieve natural and clear speech output.
[0319] "Personalized learning support" is a method of providing learning support tailored to the user's needs and level, thereby offering an efficient learning experience.
[0320] In order to implement this invention, the combination of server, user terminal, and network infrastructure is crucial. The specific configuration is described below.
[0321] First, the server analyzes digital information data using natural language processing technology. The software used here includes natural language processing libraries such as SpaCy and Transformers. This allows for the automatic identification of the difficulty level and important themes of the text, enabling information processing tailored to specific users.
[0322] The analyzed data is simplified based on the user's level of understanding. This process applies computational methods tailored to the user's characteristics, transforming complex information into a simpler form. The simplified text is then converted into audio using speech generation technology such as Google Text-to-Speech, and presented to the user.
[0323] The converted audio information and simplified text are transmitted to the user terminal via the network. The user terminal has an interface for displaying the received information and playing the audio. This allows the user to easily obtain information both visually and aurally.
[0324] For example, if a user selects a specialized academic book, the server analyzes its content, simplifies it according to the user's level of understanding, and then provides it as audio that can be listened to even while commuting. In this way, users can effectively learn while on the go.
[0325] An example of a prompt message is shown below.
[0326] Text: 'Please enter a portion of a complex book here.'
[0327] User level: 'Intermediate'
[0328] Output: 'Simplified text'
[0329] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0330] Step 1:
[0331] The user selects and inputs digital information data into the terminal, and the terminal sends that data to the server. The input here is text data from ebooks or online articles, and the output is raw text data sent to the server. The terminal uses the HTTP protocol to send the data to the server.
[0332] Step 2:
[0333] The server parses the received text data. The input is raw text data, and the output is information identifying the difficulty level and important themes of the text. The server uses natural language processing libraries (e.g., SpaCy or Transformers) to identify the structure and main topics of the text and generate metadata for the data.
[0334] Step 3:
[0335] The server simplifies text data based on identification information and user comprehension settings. The input is identification information and user comprehension settings, and the output is simplified text. The server uses specific algorithms (e.g., generative AI models) to convert complex technical terms into simpler expressions and summarize redundant parts.
[0336] Step 4:
[0337] The server converts simplified text into audio data. The input is simplified text, and the output is audio data. The server uses speech generation technology (e.g., Google Text-to-Speech) to generate the text data as an audio file.
[0338] Step 5:
[0339] The server sends the generated audio data and simplified text to the terminal. The input is the audio data and simplified text, and the output is the terminal that receives this data. Since the server streams the data over the network, the user can obtain information in real time.
[0340] Step 6:
[0341] The user plays audio and views simplified text on their device. Input consists of audio data and text sent from the server, while output is information the user hears and sees. The device provides an audio player and text viewer to support the user's learning experience.
[0342] Step 7:
[0343] Users send feedback from their devices to the server. The input is user feedback information, and the output is data stored on the server. The device provides a feedback form and reports the user's experience to the server.
[0344] Step 8:
[0345] The server analyzes the collected feedback and uses it to improve the system. The input is feedback data provided by users, and the output is an optimized system configuration. By analyzing this feedback, the server periodically improves its generation algorithm, aiming to provide more effective information.
[0346] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0347] This invention is a system that provides an interactive learning environment tailored to the user's emotional state by combining an emotion engine. The server retrieves book data selected by the user from an external database and analyzes its contents using natural language processing technology.
[0348] After analysis, the server determines the complexity and identifies key topics, then simplifies the text based on the user's reading comprehension level. At this stage, the emotion engine activates, analyzing the user's facial expressions and voice through the user's device's camera and sensors to identify emotions.
[0349] Once an emotional state is identified, the server adjusts the parameters of its simplification algorithm to change the difficulty and expression of the text according to that emotion. For example, if the server determines that the user is tired, it will provide simpler expressions and shorter sentences. In addition, the tone and speed of the synthesized speech will change according to the user's emotion. If the user is excited, a faster tone will be used; if the user is relaxed, a slower tone will be used.
[0350] The generated audio data and refined text data are sent to the user's terminal and provided in a displayable and playable format. Users can learn efficiently by reading the simplified text or listening to the audio. Furthermore, when users input feedback into their terminal after using the system, the server utilizes this feedback to improve the algorithm and the accuracy of sentiment recognition.
[0351] For example, if the emotion engine detects stress while a user is reading a scientific textbook, the server will avoid technical jargon and simplify the grammar to make the content easier to understand. Speech synthesis will also be provided in a calmer tone, allowing users to absorb knowledge while reducing stress. This enables the invention to provide a flexible and personalized learning experience.
[0352] The following describes the processing flow.
[0353] Step 1:
[0354] The user selects the book they want to read through their device. This selection information is then sent to the server.
[0355] Step 2:
[0356] The server retrieves data for the selected books from an external database via an API and temporarily stores it in text format.
[0357] Step 3:
[0358] The server analyzes the book data using natural language processing techniques. The server identifies the complexity of the text and extracts key topics and keywords.
[0359] Step 4:
[0360] The device's built-in camera and sensors capture the user's facial expressions and voice data. The device then sends this data to an emotion engine.
[0361] Step 5:
[0362] The emotion engine analyzes the user's facial expressions and voice data to identify their current emotional state. It then sends this emotional state information to the server.
[0363] Step 6:
[0364] The server applies a text simplification algorithm based on the user's reading comprehension level and emotional state. The server then generates the simplified text.
[0365] Step 7:
[0366] The server inputs simplified text into a speech synthesis engine, which then synthesizes speech data with a tone and speed that matches the user's emotions.
[0367] Step 8:
[0368] The server sends the generated audio data and simplified text to the user's device.
[0369] Step 9:
[0370] The device provides the user with an interface for playing back received audio data. The user can play the audio and visually review the simplified text.
[0371] Step 10:
[0372] The user enters feedback into the terminal after using the system. The terminal sends the collected feedback information to the server.
[0373] Step 11:
[0374] The server analyzes feedback data and uses it to improve the entire system. It enhances the accuracy of algorithms and the sentiment engine to provide a better user experience.
[0375] (Example 2)
[0376] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0377] Traditional learning systems have faced challenges in providing learning materials that adequately consider the user's emotional state, making efficient learning difficult. Furthermore, the fixed difficulty level of documents prevents flexible adaptation to the user's reading comprehension level and understanding. Additionally, there was a lack of means to continuously improve the system based on the user's learning experience.
[0378] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0379] In this invention, the server includes means for analyzing document data using natural language processing means and simplifying the text based on the user's reading comprehension level; means for analyzing the user's emotional state in real time and dynamically adjusting the content based on that analysis; and means for collecting user feedback and using it to improve the system's accuracy. This enables a flexible and personalized learning experience.
[0380] "Natural language processing methods" are technologies that analyze document data to identify the complexity and main topics of the text.
[0381] "User reading comprehension level" is an indicator that shows how well a user can understand the content of a document.
[0382] "Processing to simplify text" refers to the process of adjusting the difficulty level of a document to suit the user.
[0383] "Emotional data" refers to information such as facial expressions and voice that is acquired in order to analyze the user's emotional state.
[0384] "Dynamic adjustment methods" refer to technologies that have the ability to change the presentation of content according to the user's emotions and state.
[0385] "Speech synthesis means" refers to technology that converts text data into speech data.
[0386] A "user terminal" is an input and output interface device used by a user.
[0387] "Feedback" refers to information about the system, including opinions and evaluations collected from users.
[0388] In order to implement this invention, the server, terminal, and user interfaces must work together in coordination.
[0389] The server first receives book requests from user terminals. Book data is retrieved from an external database and analyzed using natural language processing techniques. Generative AI models such as BERT and GPT are used for this analysis to identify the complexity and main topics of the text. Based on this, the server simplifies the text to suit the user's reading level.
[0390] Meanwhile, the user's device collects facial expressions and voice data in real time through its camera and microphone. This data is sent to a server and used for emotion recognition. This process utilizes common facial recognition APIs and voice analysis APIs.
[0391] Based on emotional data, the server dynamically adjusts text and voice expressions to match the user's emotional state. For example, if the user is stressed, technical jargon is eliminated and replaced with more concise and easily understandable language. Similarly, synthesized speech is generated with a tone and speed that matches the user's emotions.
[0392] The generated audio data and adjusted text data are sent to the user's terminal and provided visually and audibly. Through this, the user can obtain an optimized learning experience. Furthermore, as the user provides feedback during the learning process, the server uses it to improve the entire system.
[0393] For example, when a user is reading a physics textbook, the server simplifies the text and adjusts the audio data based on acquired book data and real-time sentiment analysis, allowing the user to absorb the knowledge in a relaxed state.
[0394] An example of a prompt might be, "Analyze the user's facial expressions and voice to estimate their emotions and provide appropriate learning materials."
[0395] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0396] Step 1:
[0397] The server receives a book selection request sent from the user's terminal. The input is the title or identifier of the book selected by the user. The server sends a query to an external database to retrieve the book data. The retrieved book data is then output.
[0398] Step 2:
[0399] The server analyzes acquired book data using natural language processing techniques. The input is the text data of the books. Generative AI models (e.g., BERT, GPT) are used to analyze the complexity of the text and identify key topics. The output provides text feature information based on the analysis results.
[0400] Step 3:
[0401] The device captures the user's facial expressions and voice data. Inputs include real-time video and audio data. The device sends this data to the server. The server receives raw sensor data based on facial expressions and voice as output.
[0402] Step 4:
[0403] The server analyzes emotional data transmitted from the terminal. Facial and voice data are used as input. Standard facial recognition and voice analysis technologies are applied to identify the user's emotional state. Information regarding the user's emotional state is generated as output.
[0404] Step 5:
[0405] The server dynamically adjusts the text of book data based on the analyzed emotional state. Inputs are user emotional information and book feature information. A simplification algorithm is applied, eliminating technical jargon as needed and converting the text into concise language. The output is the adjusted text data.
[0406] Step 6:
[0407] The server performs speech synthesis based on simplified text. The input consists of new text data and parameters to reflect the user's emotions. This results in output speech data generated with a tone and speed appropriate to the emotional state.
[0408] Step 7:
[0409] The server sends the adjusted text data and generated audio data to the user's terminal. The input is this data, which can then be displayed and played back visually and audibly on the terminal. The output is content in a format usable by the user.
[0410] Step 8:
[0411] After completing the learning process, users input feedback into their device. This feedback includes their opinions and evaluations. The device sends this feedback to a server, and the output is data that is used to improve the system.
[0412] (Application Example 2)
[0413] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0414] Traditional learning systems have a problem in that they cannot flexibly respond to the individual emotional states of users, making it difficult to provide an appropriate learning environment. In particular, there was no mechanism to adjust the content and expression to match the user's emotions when they were feeling stressed or fatigued. As a result, users may experience decreased learning efficiency and a decline in motivation to learn.
[0415] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0416] In this invention, the server includes means for analyzing e-book data using natural language processing means to automatically identify the complexity and main topics of the text, means for analyzing the user's facial expressions and vocal state using emotion recognition means to identify emotions, and means for adjusting the difficulty level and expression of the text according to the emotional state. This makes it possible to adjust the learning content and expression to suit the user's emotional state.
[0417] "Natural language processing methods" are technologies that analyze e-book data and automatically identify the complexity and main topics of the text.
[0418] "User comprehension level" is an indicator that shows how well a user can understand a text, and it is used to adjust the learning content.
[0419] "Speech formation means" refers to a technology that generates speech information from simplified text.
[0420] A "user device" is a terminal used by a user, and is a device that enables the display and playback of audio information and simplified text.
[0421] "Emotion recognition means" refers to technology that identifies emotions by analyzing the user's facial expressions and vocal state.
[0422] "Emotional state" refers to the temporary emotional state a user is experiencing, and it is information that influences the difficulty level and expression adjustments of the learning content.
[0423] This invention realizes a system that provides an interactive learning environment that responds to the user's emotional state. First, the server retrieves e-book data from an external data store. This data is automatically analyzed using natural language processing to identify its complexity and main topics. Based on this, an algorithm is applied to simplify the text based on the user's level of understanding.
[0424] Next, the device collects data on the user's facial expressions and voice, and identifies emotions using emotion recognition tools. This process utilizes APIs such as Google Cloud Vision API and voice analysis API to obtain real-time emotion information from the user. Based on this obtained emotion information, the server adjusts the difficulty and expression of the simplified text, and generates voice information from the adjusted text using speech generation tools.
[0425] As a concrete example, if a user is reading a physics article on their smartphone and the device detects signs of fatigue, the server will respond by converting the text into simpler language, avoiding complex terminology, and generating voice information in a calmer tone. In this case, a speech synthesis API such as Amazon Polly would be used.
[0426] Finally, this adjusted text and generated audio information are sent to the user's device and played back as both text and audio. Through this, the user can receive a personalized learning experience, improving learning efficiency and motivation. The system is designed so that the generating AI model makes optimal adjustments based on a prompt message such as, "How should the content be simplified and the voice tone adjusted if the user is judged to be tired?"
[0427] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0428] Step 1:
[0429] The server retrieves e-book data from an external data store. This data is passed to a natural language processing system to analyze the complexity of the text and its main topics. The information obtained from this analysis (degree of complexity and main topics) becomes the output.
[0430] Step 2:
[0431] Based on the output of Step 1, the server applies an algorithm to simplify the text to match the user's level of understanding. The input consists of complexity information and the user's level of understanding, and the server generates a simplified text based on this. The resulting simplified document becomes the output.
[0432] Step 3:
[0433] The device collects the user's facial expressions and voice through its camera and microphone, and transmits this data to an emotion recognition system for analysis. The input is the user's facial expressions and voice data, and the emotion recognition result is output.
[0434] Step 4:
[0435] The server uses the sentiment recognition results obtained in step 3 to adjust the difficulty level and expression of the simplified text. The input is the simplified text and the sentiment recognition results, and the output is the adjusted text.
[0436] Step 5:
[0437] The server converts the adjusted text into speech data using speech synthesis means. The input is the adjusted text, and the speech data is output by speech synthesis.
[0438] Step 6:
[0439] The terminal receives audio data and edited text sent from the server, and displays and plays them for the user. The input is the data received from the server, and the output is the presentation of visual and auditory information to the user.
[0440] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0441] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0442] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0443] [Third Embodiment]
[0444] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0445] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0446] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0447] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0448] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0449] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0450] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0451] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0452] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0453] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0454] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0455] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0456] To implement this invention, a server, user terminals, and network infrastructure are primarily required. The server receives ebook data and analyzes its content using natural language processing technology. Specifically, the server analyzes the data, determines the complexity of the text, and identifies key topics and keywords.
[0457] Next, the server applies a simplification algorithm according to the reading comprehension level set by the user. This algorithm transforms the text into a simpler, more understandable format. For example, it replaces technical terms with plainer language and summarizes redundant parts.
[0458] The simplified text is then passed to a speech synthesis system. Here, the server converts the text data into synthesized speech. The synthesized speech data is configured to be clear and natural, ensuring that users can listen to it without stress.
[0459] The generated audio data and simplified text are transmitted to the user's terminal via the network. The terminal has an interface to provide the received data to the user, allowing playback, stopping, and rewinding of the audio. Users can also listen to the audio while viewing the simplified text. In this way, it becomes possible to provide efficient learning opportunities even for users who have difficulty reading printed text.
[0460] Furthermore, feedback provided by users after using the system is sent to the server via the terminal. The server analyzes the collected feedback and uses it to improve the system. This allows the service to continuously improve the user experience.
[0461] As a concrete example, suppose a user wants to read a specialized book. The server analyzes the data of that book and simplifies the content to make it easier for the user to understand. Furthermore, by providing this simplified content as audio, learning becomes possible even in situations where reading printed text is difficult, such as during commutes or while traveling. In this way, the invention realizes the provision of an efficient and high-quality learning experience to a diverse range of users.
[0462] The following describes the processing flow.
[0463] Step 1:
[0464] The server retrieves the book data specified by the user's terminal from an external database or cloud storage. The server temporarily stores the retrieved data in text format.
[0465] Step 2:
[0466] The server uses natural language processing technology to analyze the text data of books, identifying the structure and complexity of the writing. It also extracts key topics and keywords to grasp the overall content.
[0467] Step 3:
[0468] The server applies a text simplification algorithm based on the reading comprehension level set by the user. The algorithm adjusts the difficulty of vocabulary according to the settings and reconstructs sentences into a more concise form.
[0469] Step 4:
[0470] The server passes simplified text data to a speech synthesis engine to generate synthesized speech. The generated speech is natural and adjusted to be easy for the user to understand.
[0471] Step 5:
[0472] The server sends the generated audio data and simplified text data to the user's device. Encryption technology is used for data transfer for security reasons.
[0473] Step 6:
[0474] The device provides an interface for playing back received audio data. Users can use the device to control playback, pause, and rewind of the audio. It also supports visual learning by displaying simplified text.
[0475] Step 7:
[0476] The user enters feedback via a terminal after using the system. The terminal collects the user's feedback and sends it to the server.
[0477] Step 8:
[0478] The server analyzes the feedback it receives and uses it to improve the service. This will further enhance the user experience in future use.
[0479] (Example 1)
[0480] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0481] In our information-driven society, various document data are available as ebooks, but many users find it difficult to understand complex specialized and academic books. In particular, there is a problem in that users with different reading comprehension levels have difficulty obtaining information tailored to their individual needs. Furthermore, while providing audio information is important for users who have difficulty absorbing information visually, conventional technology has limitations in quality. In addition, mechanisms for efficiently utilizing user feedback to improve the system are not yet fully established.
[0482] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0483] In this invention, the server includes means for analyzing document data using an information processing device and automatically extracting the complexity of the text and the main topics; means for applying calculation means to simplify the document based on the user's reading ability; and means for generating audio information from the simplified document using a speech generation device. This enables the provision of documents tailored to different reading levels and the generation of high-quality audio information, as well as continuous improvement of the system by utilizing user feedback.
[0484] An "information processing device" is a device used to analyze data and identify and transform its contents.
[0485] "Document data" refers to data containing text information stored in electronic format.
[0486] "Complexity" refers to the difficulty of understanding the sentences within a document, or to its syntactic and lexical complexity.
[0487] "Main topic" refers to the most important or central topic or theme within the document data.
[0488] "User reading comprehension ability" refers to the degree to which individual users are able to read and understand documents.
[0489] "Computational means" refers to a method of executing algorithms or programs for performing specific processes or transformations.
[0490] A "speech generation device" is a device or system for converting text information into speech.
[0491] "Auditory information" refers to information in the form of sound that can be recognized by hearing, generated by a speech generation device.
[0492] "Feedback" refers to opinions and evaluations provided by users after using a system, and is used to improve the system.
[0493] An "external data storage device" is an external data storage facility accessed via a network, used to retrieve necessary data.
[0494] To implement this invention, a server, user terminals, and a communication network infrastructure are required. First, the server receives document data, such as ebooks, via the network. This data is provided in a common format, making it accessible to users.
[0495] Next, the server utilizes a natural language processing library (e.g., NLTK or spaCy) to analyze the received document data. Here, it identifies the complexity and main topics of the document and organizes the information. This allows the server to understand the document's structure and content, laying the foundation for providing content tailored to the user.
[0496] Subsequently, the server uses a computational method (algorithm) to generate a simplified document based on the reading comprehension level set by the user. The algorithm replaces complex technical terms with simpler language and summarizes redundant content. Techniques used in this process include document summarization algorithms and paraphrasing techniques.
[0497] The simplified text is then converted into audio information via a speech generator. Here, a speech synthesis API (e.g., Google Text-to-Speech API or Amazon Polly) is used to generate natural and clear synthesized speech. This speech is particularly effective as a visually independent means of communication.
[0498] The generated audio information and simplified text are transmitted to the user's terminal via the network. The user's terminal receives this information, displays it for easy user access, and provides audio playback, pause, and rewind functions. For example, a user can listen to the audio information using earphones while commuting and simultaneously review the simplified text.
[0499] As a concrete example, a user who wants to understand the textbook "Digital Signal Processing" requests the data from the server. The server converts the "complex mathematical explanations" into "basic data analysis methods" and provides them to the user as audio. An example of a prompt used in this process would be, "Please convert this text into simple language that even a middle school student can understand."
[0500] In this way, the server and user terminals work together to create a system that provides efficient and easy-to-understand information.
[0501] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0502] Step 1:
[0503] The server receives ebook data via the network based on user requests. It takes ebook file format data (e.g., EPUB, PDF) as input and outputs parseable text data. Specifically, the server downloads data from specified resources using network protocols.
[0504] Step 2:
[0505] The server analyzes received text data using natural language processing libraries (e.g., NLTK, spaCy). It takes text data as input and outputs metadata identifying sentence complexity and main topics. The server performs data processing such as tokenization, part-of-speech tagging, and noun phrase extraction. Specifically, the server executes a parsing algorithm to extract important keywords and concepts from the text.
[0506] Step 3:
[0507] The server simplifies text, taking into account the user's reading level. It takes parsed metadata and user reading level information as input and outputs simplified text. The server applies algorithms to convert complex terminology into simpler expressions and summarize redundant sentences. Specifically, the server uses parsing to reconstruct sentences into a more concise form.
[0508] Step 4:
[0509] The server converts simplified text into speech information using a speech generator. Simplified text is the input, and synthesized speech data is the output. The server uses a speech synthesis API (e.g., Google Text-to-Speech API) to convert the text into an audio file (e.g., MP3 format). Specifically, the server adjusts phonemes and intonation to generate natural and clear speech.
[0510] Step 5:
[0511] The server sends the generated audio information and simplified text to the user terminal over the network. The input consists of an audio file and simplified text, and the output is sent in a format that can be displayed and played on the user terminal. The server composes the data package and prepares it for transmission according to the transmission protocol. Specifically, the server binds to the destination address and transmits the data through the specified port.
[0512] Step 6:
[0513] The user terminal provides received audio and text information, enabling the user to listen to and view it. It receives data from the server as input and outputs an interactive display for the user. The terminal launches an audio player and text viewer, providing playback, pause, and rewind functions. Specifically, the terminal displays play and pause buttons on the user interface and outputs audio.
[0514] Step 7:
[0515] Users provide feedback after using the system. The input consists of the user's subjective evaluation and comments, and the output is feedback data sent to the server. Users submit their opinions by entering their evaluation through a feedback form on their terminal and pressing the submit button. Specifically, users can select evaluation items and leave comments in the text field.
[0516] Step 8:
[0517] The server analyzes the collected feedback to help improve the system. It receives feedback data from users as input and outputs information for developing system improvement strategies. The server uses text analysis techniques to extract common problems and improvement requests from the feedback. Specifically, the server stores the feedback data in a database and uses machine learning algorithms to analyze patterns.
[0518] (Application Example 1)
[0519] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0520] In recent years, with the spread of e-books and digital information, there has been a growing demand for information to be acquired efficiently and in an easily understandable format by users at various levels. However, much of this information is specialized and complex, making it difficult for some users to comprehend. Furthermore, there is a need to efficiently acquire information even in situations where directly reading text is difficult, such as during commutes or while traveling. These problems are particularly pronounced when using specialized books or materials containing advanced content. Therefore, there is a need for systems that simplify e-books and digital information according to the user's level of understanding and provide it in audio format.
[0521] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0522] In this invention, the server includes means for analyzing digital information data using natural language processing technology to automatically identify the difficulty level and important themes of the text, means for applying computational means to simplify the text according to the user's level of understanding, and means for creating audio information from the simplified text using speech generation technology. This enables personalized learning assistance for the user by simplifying the text and converting it to speech based on the digital information content selected by the user.
[0523] "Natural language processing technology" refers to the technology that enables computers to understand, interpret, and generate natural language, and is used to analyze the meaning of text.
[0524] "Digital information data" refers to a collection of information stored in digital format, including ebooks and online articles.
[0525] "Difficulty level of a text" is an indicator of the level of comprehension required for a text, and is evaluated based on factors such as the frequency of use of complex grammar and technical terms.
[0526] An "important theme" refers to a particularly important or central topic within a text or piece of information, requiring a deep understanding of its content.
[0527] "Comprehension level" is an indicator of how accurately a user can understand information, and it depends on the individual's knowledge level and experience.
[0528] "Computational means" include mathematical methods and algorithms for simplifying complex information according to the characteristics of the user.
[0529] "Speech generation technology" is a technology that converts text data into speech data and is used to achieve natural and clear speech output.
[0530] "Personalized learning support" is a method of providing learning support tailored to the user's needs and level, thereby offering an efficient learning experience.
[0531] In order to implement this invention, the combination of server, user terminal, and network infrastructure is crucial. The specific configuration is described below.
[0532] First, the server analyzes digital information data using natural language processing technology. The software used here includes natural language processing libraries such as SpaCy and Transformers. This allows for the automatic identification of the difficulty level and important themes of the text, enabling information processing tailored to specific users.
[0533] The analyzed data is simplified based on the user's level of understanding. This process applies computational methods tailored to the user's characteristics, transforming complex information into a simpler form. The simplified text is then converted into audio using speech generation technology such as Google Text-to-Speech, and presented to the user.
[0534] The converted audio information and simplified text are transmitted to the user terminal via the network. The user terminal has an interface for displaying the received information and playing the audio. This allows the user to easily obtain information both visually and aurally.
[0535] For example, if a user selects a specialized academic book, the server analyzes its content, simplifies it according to the user's level of understanding, and then provides it as audio that can be listened to even while commuting. In this way, users can effectively learn while on the go.
[0536] An example of a prompt message is shown below.
[0537] Text: 'Please enter a portion of a complex book here.'
[0538] User level: 'Intermediate'
[0539] Output: 'Simplified text'
[0540] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0541] Step 1:
[0542] The user selects and inputs digital information data into the terminal, and the terminal sends that data to the server. The input here is text data from ebooks or online articles, and the output is raw text data sent to the server. The terminal uses the HTTP protocol to send the data to the server.
[0543] Step 2:
[0544] The server parses the received text data. The input is raw text data, and the output is information identifying the difficulty level and important themes of the text. The server uses natural language processing libraries (e.g., SpaCy or Transformers) to identify the structure and main topics of the text and generate metadata for the data.
[0545] Step 3:
[0546] The server simplifies text data based on identification information and user comprehension settings. The input is identification information and user comprehension settings, and the output is simplified text. The server uses specific algorithms (e.g., generative AI models) to convert complex technical terms into simpler expressions and summarize redundant parts.
[0547] Step 4:
[0548] The server converts simplified text into audio data. The input is simplified text, and the output is audio data. The server uses speech generation technology (e.g., Google Text-to-Speech) to generate the text data as an audio file.
[0549] Step 5:
[0550] The server sends the generated audio data and simplified text to the terminal. The input is the audio data and simplified text, and the output is the terminal that receives this data. Since the server streams the data over the network, the user can obtain information in real time.
[0551] Step 6:
[0552] The user plays audio and views simplified text on their device. Input consists of audio data and text sent from the server, while output is information the user hears and sees. The device provides an audio player and text viewer to support the user's learning experience.
[0553] Step 7:
[0554] Users send feedback from their devices to the server. The input is user feedback information, and the output is data stored on the server. The device provides a feedback form and reports the user's experience to the server.
[0555] Step 8:
[0556] The server analyzes the collected feedback and uses it to improve the system. The input is feedback data provided by users, and the output is an optimized system configuration. By analyzing this feedback, the server periodically improves its generation algorithm, aiming to provide more effective information.
[0557] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0558] This invention is a system that provides an interactive learning environment tailored to the user's emotional state by combining an emotion engine. The server retrieves book data selected by the user from an external database and analyzes its contents using natural language processing technology.
[0559] After analysis, the server determines the complexity and identifies key topics, then simplifies the text based on the user's reading comprehension level. At this stage, the emotion engine activates, analyzing the user's facial expressions and voice through the user's device's camera and sensors to identify emotions.
[0560] Once an emotional state is identified, the server adjusts the parameters of its simplification algorithm to change the difficulty and expression of the text according to that emotion. For example, if the server determines that the user is tired, it will provide simpler expressions and shorter sentences. In addition, the tone and speed of the synthesized speech will change according to the user's emotion. If the user is excited, a faster tone will be used; if the user is relaxed, a slower tone will be used.
[0561] The generated audio data and refined text data are sent to the user's terminal and provided in a displayable and playable format. Users can learn efficiently by reading the simplified text or listening to the audio. Furthermore, when users input feedback into their terminal after using the system, the server utilizes this feedback to improve the algorithm and the accuracy of sentiment recognition.
[0562] For example, if the emotion engine detects stress while a user is reading a scientific textbook, the server will avoid technical jargon and simplify the grammar to make the content easier to understand. Speech synthesis will also be provided in a calmer tone, allowing users to absorb knowledge while reducing stress. This enables the invention to provide a flexible and personalized learning experience.
[0563] The following describes the processing flow.
[0564] Step 1:
[0565] The user selects the book they want to read through their device. This selection information is then sent to the server.
[0566] Step 2:
[0567] The server retrieves data for the selected books from an external database via an API and temporarily stores it in text format.
[0568] Step 3:
[0569] The server analyzes the book data using natural language processing techniques. The server identifies the complexity of the text and extracts key topics and keywords.
[0570] Step 4:
[0571] The device's built-in camera and sensors capture the user's facial expressions and voice data. The device then sends this data to an emotion engine.
[0572] Step 5:
[0573] The emotion engine analyzes the user's facial expressions and voice data to identify their current emotional state. It then sends this emotional state information to the server.
[0574] Step 6:
[0575] The server applies a text simplification algorithm based on the user's reading comprehension level and emotional state. The server then generates the simplified text.
[0576] Step 7:
[0577] The server inputs simplified text into a speech synthesis engine, which then synthesizes speech data with a tone and speed that matches the user's emotions.
[0578] Step 8:
[0579] The server sends the generated audio data and simplified text to the user's device.
[0580] Step 9:
[0581] The device provides the user with an interface for playing back received audio data. The user can play the audio and visually review the simplified text.
[0582] Step 10:
[0583] The user enters feedback into the terminal after using the system. The terminal sends the collected feedback information to the server.
[0584] Step 11:
[0585] The server analyzes feedback data and uses it to improve the entire system. It enhances the accuracy of algorithms and the sentiment engine to provide a better user experience.
[0586] (Example 2)
[0587] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0588] Traditional learning systems have faced challenges in providing learning materials that adequately consider the user's emotional state, making efficient learning difficult. Furthermore, the fixed difficulty level of documents prevents flexible adaptation to the user's reading comprehension level and understanding. Additionally, there was a lack of means to continuously improve the system based on the user's learning experience.
[0589] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0590] In this invention, the server includes means for analyzing document data using natural language processing means and simplifying the text based on the user's reading comprehension level; means for analyzing the user's emotional state in real time and dynamically adjusting the content based on that analysis; and means for collecting user feedback and using it to improve the system's accuracy. This enables a flexible and personalized learning experience.
[0591] "Natural language processing methods" are technologies that analyze document data to identify the complexity and main topics of the text.
[0592] "User reading comprehension level" is an indicator that shows how well a user can understand the content of a document.
[0593] "Processing to simplify text" refers to the process of adjusting the difficulty level of a document to suit the user.
[0594] "Emotional data" refers to information such as facial expressions and voice that is acquired in order to analyze the user's emotional state.
[0595] "Dynamic adjustment methods" refer to technologies that have the ability to change the presentation of content according to the user's emotions and state.
[0596] "Speech synthesis means" refers to technology that converts text data into speech data.
[0597] A "user terminal" is an input and output interface device used by a user.
[0598] "Feedback" refers to information about the system, including opinions and evaluations collected from users.
[0599] In order to implement this invention, the server, terminal, and user interfaces must work together in coordination.
[0600] The server first receives book requests from user terminals. Book data is retrieved from an external database and analyzed using natural language processing techniques. Generative AI models such as BERT and GPT are used for this analysis to identify the complexity and main topics of the text. Based on this, the server simplifies the text to suit the user's reading level.
[0601] Meanwhile, the user's device collects facial expressions and voice data in real time through its camera and microphone. This data is sent to a server and used for emotion recognition. This process utilizes common facial recognition APIs and voice analysis APIs.
[0602] Based on emotional data, the server dynamically adjusts text and voice expressions to match the user's emotional state. For example, if the user is stressed, technical jargon is eliminated and replaced with more concise and easily understandable language. Similarly, synthesized speech is generated with a tone and speed that matches the user's emotions.
[0603] The generated audio data and adjusted text data are sent to the user's terminal and provided visually and audibly. Through this, the user can obtain an optimized learning experience. Furthermore, as the user provides feedback during the learning process, the server uses it to improve the entire system.
[0604] For example, when a user is reading a physics textbook, the server simplifies the text and adjusts the audio data based on acquired book data and real-time sentiment analysis, allowing the user to absorb the knowledge in a relaxed state.
[0605] An example of a prompt might be, "Analyze the user's facial expressions and voice to estimate their emotions and provide appropriate learning materials."
[0606] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0607] Step 1:
[0608] The server receives a book selection request sent from the user's terminal. The input is the title or identifier of the book selected by the user. The server sends a query to an external database to retrieve the book data. The retrieved book data is then output.
[0609] Step 2:
[0610] The server analyzes acquired book data using natural language processing techniques. The input is the text data of the books. Generative AI models (e.g., BERT, GPT) are used to analyze the complexity of the text and identify key topics. The output provides text feature information based on the analysis results.
[0611] Step 3:
[0612] The device captures the user's facial expressions and voice data. Inputs include real-time video and audio data. The device sends this data to the server. The server receives raw sensor data based on facial expressions and voice as output.
[0613] Step 4:
[0614] The server analyzes emotional data transmitted from the terminal. Facial and voice data are used as input. Standard facial recognition and voice analysis technologies are applied to identify the user's emotional state. Information regarding the user's emotional state is generated as output.
[0615] Step 5:
[0616] The server dynamically adjusts the text of book data based on the analyzed emotional state. Inputs are user emotional information and book feature information. A simplification algorithm is applied, eliminating technical jargon as needed and converting the text into concise language. The output is the adjusted text data.
[0617] Step 6:
[0618] The server performs speech synthesis based on simplified text. The input consists of new text data and parameters to reflect the user's emotions. This results in output speech data generated with a tone and speed appropriate to the emotional state.
[0619] Step 7:
[0620] The server sends the adjusted text data and generated audio data to the user's terminal. The input is this data, which can then be displayed and played back visually and audibly on the terminal. The output is content in a format usable by the user.
[0621] Step 8:
[0622] After completing the learning process, users input feedback into their device. This feedback includes their opinions and evaluations. The device sends this feedback to a server, and the output is data that is used to improve the system.
[0623] (Application Example 2)
[0624] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0625] Traditional learning systems have a problem in that they cannot flexibly respond to the individual emotional states of users, making it difficult to provide an appropriate learning environment. In particular, there was no mechanism to adjust the content and expression to match the user's emotions when they were feeling stressed or fatigued. As a result, users may experience decreased learning efficiency and a decline in motivation to learn.
[0626] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0627] In this invention, the server includes means for analyzing e-book data using natural language processing means to automatically identify the complexity and main topics of the text, means for analyzing the user's facial expressions and vocal state using emotion recognition means to identify emotions, and means for adjusting the difficulty level and expression of the text according to the emotional state. This makes it possible to adjust the learning content and expression to suit the user's emotional state.
[0628] "Natural language processing methods" are technologies that analyze e-book data and automatically identify the complexity and main topics of the text.
[0629] "User comprehension level" is an indicator that shows how well a user can understand a text, and it is used to adjust the learning content.
[0630] "Speech formation means" refers to a technology that generates speech information from simplified text.
[0631] A "user device" is a terminal used by a user, and is a device that enables the display and playback of audio information and simplified text.
[0632] "Emotion recognition means" refers to technology that identifies emotions by analyzing the user's facial expressions and vocal state.
[0633] "Emotional state" refers to the temporary emotional state a user is experiencing, and it is information that influences the difficulty level and expression adjustments of the learning content.
[0634] This invention realizes a system that provides an interactive learning environment that responds to the user's emotional state. First, the server retrieves e-book data from an external data store. This data is automatically analyzed using natural language processing to identify its complexity and main topics. Based on this, an algorithm is applied to simplify the text based on the user's level of understanding.
[0635] Next, the device collects data on the user's facial expressions and voice, and identifies emotions using emotion recognition tools. This process utilizes APIs such as Google Cloud Vision API and voice analysis API to obtain real-time emotion information from the user. Based on this obtained emotion information, the server adjusts the difficulty and expression of the simplified text, and generates voice information from the adjusted text using speech generation tools.
[0636] As a concrete example, if a user is reading a physics article on their smartphone and the device detects signs of fatigue, the server will respond by converting the text into simpler language, avoiding complex terminology, and generating voice information in a calmer tone. In this case, a speech synthesis API such as Amazon Polly would be used.
[0637] Finally, this adjusted text and generated audio information are sent to the user's device and played back as both text and audio. Through this, the user can receive a personalized learning experience, improving learning efficiency and motivation. The system is designed so that the generating AI model makes optimal adjustments based on a prompt message such as, "How should the content be simplified and the voice tone adjusted if the user is judged to be tired?"
[0638] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0639] Step 1:
[0640] The server retrieves e-book data from an external data store. This data is passed to a natural language processing system to analyze the complexity of the text and its main topics. The information obtained from this analysis (degree of complexity and main topics) becomes the output.
[0641] Step 2:
[0642] Based on the output of Step 1, the server applies an algorithm to simplify the text to match the user's level of understanding. The input consists of complexity information and the user's level of understanding, and the server generates a simplified text based on this. The resulting simplified document becomes the output.
[0643] Step 3:
[0644] The device collects the user's facial expressions and voice through its camera and microphone, and transmits this data to an emotion recognition system for analysis. The input is the user's facial expressions and voice data, and the emotion recognition result is output.
[0645] Step 4:
[0646] The server uses the sentiment recognition results obtained in step 3 to adjust the difficulty level and expression of the simplified text. The input is the simplified text and the sentiment recognition results, and the output is the adjusted text.
[0647] Step 5:
[0648] The server converts the adjusted text into speech data using speech synthesis means. The input is the adjusted text, and the speech data is output by speech synthesis.
[0649] Step 6:
[0650] The terminal receives audio data and edited text sent from the server, and displays and plays them for the user. The input is the data received from the server, and the output is the presentation of visual and auditory information to the user.
[0651] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0652] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0653] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0654] [Fourth Embodiment]
[0655] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0656] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0657] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0658] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0659] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0660] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0661] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0662] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0663] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0664] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0665] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0666] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0667] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0668] To implement this invention, a server, user terminals, and network infrastructure are primarily required. The server receives ebook data and analyzes its content using natural language processing technology. Specifically, the server analyzes the data, determines the complexity of the text, and identifies key topics and keywords.
[0669] Next, the server applies a simplification algorithm according to the reading comprehension level set by the user. This algorithm transforms the text into a simpler, more understandable format. For example, it replaces technical terms with plainer language and summarizes redundant parts.
[0670] The simplified text is then passed to a speech synthesis system. Here, the server converts the text data into synthesized speech. The synthesized speech data is configured to be clear and natural, ensuring that users can listen to it without stress.
[0671] The generated audio data and simplified text are transmitted to the user's terminal via the network. The terminal has an interface to provide the received data to the user, allowing playback, stopping, and rewinding of the audio. Users can also listen to the audio while viewing the simplified text. In this way, it becomes possible to provide efficient learning opportunities even for users who have difficulty reading printed text.
[0672] Furthermore, feedback provided by users after using the system is sent to the server via the terminal. The server analyzes the collected feedback and uses it to improve the system. This allows the service to continuously improve the user experience.
[0673] As a concrete example, suppose a user wants to read a specialized book. The server analyzes the data of that book and simplifies the content to make it easier for the user to understand. Furthermore, by providing this simplified content as audio, learning becomes possible even in situations where reading printed text is difficult, such as during commutes or while traveling. In this way, the invention realizes the provision of an efficient and high-quality learning experience to a diverse range of users.
[0674] The following describes the processing flow.
[0675] Step 1:
[0676] The server retrieves the book data specified by the user's terminal from an external database or cloud storage. The server temporarily stores the retrieved data in text format.
[0677] Step 2:
[0678] The server uses natural language processing technology to analyze the text data of books, identifying the structure and complexity of the writing. It also extracts key topics and keywords to grasp the overall content.
[0679] Step 3:
[0680] The server applies a text simplification algorithm based on the reading comprehension level set by the user. The algorithm adjusts the difficulty of vocabulary according to the settings and reconstructs sentences into a more concise form.
[0681] Step 4:
[0682] The server passes simplified text data to a speech synthesis engine to generate synthesized speech. The generated speech is natural and adjusted to be easy for the user to understand.
[0683] Step 5:
[0684] The server sends the generated audio data and simplified text data to the user's device. Encryption technology is used for data transfer for security reasons.
[0685] Step 6:
[0686] The device provides an interface for playing back received audio data. Users can use the device to control playback, pause, and rewind of the audio. It also supports visual learning by displaying simplified text.
[0687] Step 7:
[0688] The user enters feedback via a terminal after using the system. The terminal collects the user's feedback and sends it to the server.
[0689] Step 8:
[0690] The server analyzes the feedback it receives and uses it to improve the service. This will further enhance the user experience in future use.
[0691] (Example 1)
[0692] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0693] In our information-driven society, various document data are available as ebooks, but many users find it difficult to understand complex specialized and academic books. In particular, there is a problem in that users with different reading comprehension levels have difficulty obtaining information tailored to their individual needs. Furthermore, while providing audio information is important for users who have difficulty absorbing information visually, conventional technology has limitations in quality. In addition, mechanisms for efficiently utilizing user feedback to improve the system are not yet fully established.
[0694] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0695] In this invention, the server includes means for analyzing document data using an information processing device and automatically extracting the complexity of the text and the main topics; means for applying calculation means to simplify the document based on the user's reading ability; and means for generating audio information from the simplified document using a speech generation device. This enables the provision of documents tailored to different reading levels and the generation of high-quality audio information, as well as continuous improvement of the system by utilizing user feedback.
[0696] An "information processing device" is a device used to analyze data and identify and transform its contents.
[0697] "Document data" refers to data containing text information stored in electronic format.
[0698] "Complexity" refers to the difficulty of understanding the sentences within a document, or to its syntactic and lexical complexity.
[0699] "Main topic" refers to the most important or central topic or theme within the document data.
[0700] "User reading comprehension ability" refers to the degree to which individual users are able to read and understand documents.
[0701] "Computational means" refers to a method of executing algorithms or programs for performing specific processes or transformations.
[0702] A "speech generation device" is a device or system for converting text information into speech.
[0703] "Auditory information" refers to information in the form of sound that can be recognized by hearing, generated by a speech generation device.
[0704] "Feedback" refers to opinions and evaluations provided by users after using a system, and is used to improve the system.
[0705] An "external data storage device" is an external data storage facility accessed via a network, used to retrieve necessary data.
[0706] To implement this invention, a server, user terminals, and a communication network infrastructure are required. First, the server receives document data, such as ebooks, via the network. This data is provided in a common format, making it accessible to users.
[0707] Next, the server utilizes a natural language processing library (e.g., NLTK or spaCy) to analyze the received document data. Here, it identifies the complexity and main topics of the document and organizes the information. This allows the server to understand the document's structure and content, laying the foundation for providing content tailored to the user.
[0708] Subsequently, the server uses a computational method (algorithm) to generate a simplified document based on the reading comprehension level set by the user. The algorithm replaces complex technical terms with simpler language and summarizes redundant content. Techniques used in this process include document summarization algorithms and paraphrasing techniques.
[0709] The simplified text is then converted into audio information via a speech generator. Here, a speech synthesis API (e.g., Google Text-to-Speech API or Amazon Polly) is used to generate natural and clear synthesized speech. This speech is particularly effective as a visually independent means of communication.
[0710] The generated audio information and simplified text are transmitted to the user's terminal via the network. The user's terminal receives this information, displays it for easy user access, and provides audio playback, pause, and rewind functions. For example, a user can listen to the audio information using earphones while commuting and simultaneously review the simplified text.
[0711] As a concrete example, a user who wants to understand the textbook "Digital Signal Processing" requests the data from the server. The server converts the "complex mathematical explanations" into "basic data analysis methods" and provides them to the user as audio. An example of a prompt used in this process would be, "Please convert this text into simple language that even a middle school student can understand."
[0712] In this way, the server and user terminals work together to create a system that provides efficient and easy-to-understand information.
[0713] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0714] Step 1:
[0715] The server receives ebook data via the network based on user requests. It takes ebook file format data (e.g., EPUB, PDF) as input and outputs parseable text data. Specifically, the server downloads data from specified resources using network protocols.
[0716] Step 2:
[0717] The server analyzes received text data using natural language processing libraries (e.g., NLTK, spaCy). It takes text data as input and outputs metadata identifying sentence complexity and main topics. The server performs data processing such as tokenization, part-of-speech tagging, and noun phrase extraction. Specifically, the server executes a parsing algorithm to extract important keywords and concepts from the text.
[0718] Step 3:
[0719] The server simplifies text, taking into account the user's reading level. It takes parsed metadata and user reading level information as input and outputs simplified text. The server applies algorithms to convert complex terminology into simpler expressions and summarize redundant sentences. Specifically, the server uses parsing to reconstruct sentences into a more concise form.
[0720] Step 4:
[0721] The server converts simplified text into speech information using a speech generator. Simplified text is the input, and synthesized speech data is the output. The server uses a speech synthesis API (e.g., Google Text-to-Speech API) to convert the text into an audio file (e.g., MP3 format). Specifically, the server adjusts phonemes and intonation to generate natural and clear speech.
[0722] Step 5:
[0723] The server sends the generated audio information and simplified text to the user terminal over the network. The input consists of an audio file and simplified text, and the output is sent in a format that can be displayed and played on the user terminal. The server composes the data package and prepares it for transmission according to the transmission protocol. Specifically, the server binds to the destination address and transmits the data through the specified port.
[0724] Step 6:
[0725] The user terminal provides received audio and text information, enabling the user to listen to and view it. It receives data from the server as input and outputs an interactive display for the user. The terminal launches an audio player and text viewer, providing playback, pause, and rewind functions. Specifically, the terminal displays play and pause buttons on the user interface and outputs audio.
[0726] Step 7:
[0727] Users provide feedback after using the system. The input consists of the user's subjective evaluation and comments, and the output is feedback data sent to the server. Users submit their opinions by entering their evaluation through a feedback form on their terminal and pressing the submit button. Specifically, users can select evaluation items and leave comments in the text field.
[0728] Step 8:
[0729] The server analyzes the collected feedback to help improve the system. It receives feedback data from users as input and outputs information for developing system improvement strategies. The server uses text analysis techniques to extract common problems and improvement requests from the feedback. Specifically, the server stores the feedback data in a database and uses machine learning algorithms to analyze patterns.
[0730] (Application Example 1)
[0731] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0732] In recent years, with the spread of e-books and digital information, there has been a growing demand for information to be acquired efficiently and in an easily understandable format by users at various levels. However, much of this information is specialized and complex, making it difficult for some users to comprehend. Furthermore, there is a need to efficiently acquire information even in situations where directly reading text is difficult, such as during commutes or while traveling. These problems are particularly pronounced when using specialized books or materials containing advanced content. Therefore, there is a need for systems that simplify e-books and digital information according to the user's level of understanding and provide it in audio format.
[0733] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0734] In this invention, the server includes means for analyzing digital information data using natural language processing technology to automatically identify the difficulty level and important themes of the text, means for applying computational means to simplify the text according to the user's level of understanding, and means for creating audio information from the simplified text using speech generation technology. This enables personalized learning assistance for the user by simplifying the text and converting it to speech based on the digital information content selected by the user.
[0735] "Natural language processing technology" refers to the technology that enables computers to understand, interpret, and generate natural language, and is used to analyze the meaning of text.
[0736] "Digital information data" refers to a collection of information stored in digital format, including ebooks and online articles.
[0737] "Difficulty level of a text" is an indicator of the level of comprehension required for a text, and is evaluated based on factors such as the frequency of use of complex grammar and technical terms.
[0738] An "important theme" refers to a particularly important or central topic within a text or piece of information, requiring a deep understanding of its content.
[0739] "Comprehension level" is an indicator of how accurately a user can understand information, and it depends on the individual's knowledge level and experience.
[0740] "Computational means" include mathematical methods and algorithms for simplifying complex information according to the characteristics of the user.
[0741] "Speech generation technology" is a technology that converts text data into speech data and is used to achieve natural and clear speech output.
[0742] "Personalized learning support" is a method of providing learning support tailored to the user's needs and level, thereby offering an efficient learning experience.
[0743] In order to implement this invention, the combination of server, user terminal, and network infrastructure is crucial. The specific configuration is described below.
[0744] First, the server analyzes digital information data using natural language processing technology. The software used here includes natural language processing libraries such as SpaCy and Transformers. This allows for the automatic identification of the difficulty level and important themes of the text, enabling information processing tailored to specific users.
[0745] The analyzed data is simplified based on the user's level of understanding. This process applies computational methods tailored to the user's characteristics, transforming complex information into a simpler form. The simplified text is then converted into audio using speech generation technology such as Google Text-to-Speech, and presented to the user.
[0746] The converted audio information and simplified text are transmitted to the user terminal via the network. The user terminal has an interface for displaying the received information and playing the audio. This allows the user to easily obtain information both visually and aurally.
[0747] For example, if a user selects a specialized academic book, the server analyzes its content, simplifies it according to the user's level of understanding, and then provides it as audio that can be listened to even while commuting. In this way, users can effectively learn while on the go.
[0748] An example of a prompt message is shown below.
[0749] Text: 'Please enter a portion of a complex book here.'
[0750] User level: 'Intermediate'
[0751] Output: 'Simplified text'
[0752] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0753] Step 1:
[0754] The user selects and inputs digital information data into the terminal, and the terminal sends that data to the server. The input here is text data from ebooks or online articles, and the output is raw text data sent to the server. The terminal uses the HTTP protocol to send the data to the server.
[0755] Step 2:
[0756] The server parses the received text data. The input is raw text data, and the output is information identifying the difficulty level and important themes of the text. The server uses natural language processing libraries (e.g., SpaCy or Transformers) to identify the structure and main topics of the text and generate metadata for the data.
[0757] Step 3:
[0758] The server simplifies text data based on identification information and user comprehension settings. The input is identification information and user comprehension settings, and the output is simplified text. The server uses specific algorithms (e.g., generative AI models) to convert complex technical terms into simpler expressions and summarize redundant parts.
[0759] Step 4:
[0760] The server converts simplified text into audio data. The input is simplified text, and the output is audio data. The server uses speech generation technology (e.g., Google Text-to-Speech) to generate the text data as an audio file.
[0761] Step 5:
[0762] The server sends the generated audio data and simplified text to the terminal. The input is the audio data and simplified text, and the output is the terminal that receives this data. Since the server streams the data over the network, the user can obtain information in real time.
[0763] Step 6:
[0764] The user plays audio and views simplified text on their device. Input consists of audio data and text sent from the server, while output is information the user hears and sees. The device provides an audio player and text viewer to support the user's learning experience.
[0765] Step 7:
[0766] Users send feedback from their devices to the server. The input is user feedback information, and the output is data stored on the server. The device provides a feedback form and reports the user's experience to the server.
[0767] Step 8:
[0768] The server analyzes the collected feedback and uses it to improve the system. The input is feedback data provided by users, and the output is an optimized system configuration. By analyzing this feedback, the server periodically improves its generation algorithm, aiming to provide more effective information.
[0769] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0770] This invention is a system that provides an interactive learning environment tailored to the user's emotional state by combining an emotion engine. The server retrieves book data selected by the user from an external database and analyzes its contents using natural language processing technology.
[0771] After analysis, the server determines the complexity and identifies key topics, then simplifies the text based on the user's reading comprehension level. At this stage, the emotion engine activates, analyzing the user's facial expressions and voice through the user's device's camera and sensors to identify emotions.
[0772] Once an emotional state is identified, the server adjusts the parameters of its simplification algorithm to change the difficulty and expression of the text according to that emotion. For example, if the server determines that the user is tired, it will provide simpler expressions and shorter sentences. In addition, the tone and speed of the synthesized speech will change according to the user's emotion. If the user is excited, a faster tone will be used; if the user is relaxed, a slower tone will be used.
[0773] The generated audio data and refined text data are sent to the user's terminal and provided in a displayable and playable format. Users can learn efficiently by reading the simplified text or listening to the audio. Furthermore, when users input feedback into their terminal after using the system, the server utilizes this feedback to improve the algorithm and the accuracy of sentiment recognition.
[0774] For example, if the emotion engine detects stress while a user is reading a scientific textbook, the server will avoid technical jargon and simplify the grammar to make the content easier to understand. Speech synthesis will also be provided in a calmer tone, allowing users to absorb knowledge while reducing stress. This enables the invention to provide a flexible and personalized learning experience.
[0775] The following describes the processing flow.
[0776] Step 1:
[0777] The user selects the book they want to read through their device. This selection information is then sent to the server.
[0778] Step 2:
[0779] The server retrieves data for the selected books from an external database via an API and temporarily stores it in text format.
[0780] Step 3:
[0781] The server analyzes the book data using natural language processing techniques. The server identifies the complexity of the text and extracts key topics and keywords.
[0782] Step 4:
[0783] The device's built-in camera and sensors capture the user's facial expressions and voice data. The device then sends this data to an emotion engine.
[0784] Step 5:
[0785] The emotion engine analyzes the user's facial expressions and voice data to identify their current emotional state. It then sends this emotional state information to the server.
[0786] Step 6:
[0787] The server applies a text simplification algorithm based on the user's reading comprehension level and emotional state. The server then generates the simplified text.
[0788] Step 7:
[0789] The server inputs simplified text into a speech synthesis engine, which then synthesizes speech data with a tone and speed that matches the user's emotions.
[0790] Step 8:
[0791] The server sends the generated audio data and simplified text to the user's device.
[0792] Step 9:
[0793] The device provides the user with an interface for playing back received audio data. The user can play the audio and visually review the simplified text.
[0794] Step 10:
[0795] The user enters feedback into the terminal after using the system. The terminal sends the collected feedback information to the server.
[0796] Step 11:
[0797] The server analyzes feedback data and uses it to improve the entire system. It enhances the accuracy of algorithms and the sentiment engine to provide a better user experience.
[0798] (Example 2)
[0799] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0800] Traditional learning systems have faced challenges in providing learning materials that adequately consider the user's emotional state, making efficient learning difficult. Furthermore, the fixed difficulty level of documents prevents flexible adaptation to the user's reading comprehension level and understanding. Additionally, there was a lack of means to continuously improve the system based on the user's learning experience.
[0801] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0802] In this invention, the server includes means for analyzing document data using natural language processing means and simplifying the text based on the user's reading comprehension level; means for analyzing the user's emotional state in real time and dynamically adjusting the content based on that analysis; and means for collecting user feedback and using it to improve the system's accuracy. This enables a flexible and personalized learning experience.
[0803] "Natural language processing methods" are technologies that analyze document data to identify the complexity and main topics of the text.
[0804] "User reading comprehension level" is an indicator that shows how well a user can understand the content of a document.
[0805] "Processing to simplify text" refers to the process of adjusting the difficulty level of a document to suit the user.
[0806] "Emotional data" refers to information such as facial expressions and voice that is acquired in order to analyze the user's emotional state.
[0807] "Dynamic adjustment methods" refer to technologies that have the ability to change the presentation of content according to the user's emotions and state.
[0808] "Speech synthesis means" refers to technology that converts text data into speech data.
[0809] A "user terminal" is an input and output interface device used by a user.
[0810] "Feedback" refers to information about the system, including opinions and evaluations collected from users.
[0811] In order to implement this invention, the server, terminal, and user interfaces must work together in coordination.
[0812] The server first receives book requests from user terminals. Book data is retrieved from an external database and analyzed using natural language processing techniques. Generative AI models such as BERT and GPT are used for this analysis to identify the complexity and main topics of the text. Based on this, the server simplifies the text to suit the user's reading level.
[0813] Meanwhile, the user's device collects facial expressions and voice data in real time through its camera and microphone. This data is sent to a server and used for emotion recognition. This process utilizes common facial recognition APIs and voice analysis APIs.
[0814] Based on emotional data, the server dynamically adjusts text and voice expressions to match the user's emotional state. For example, if the user is stressed, technical jargon is eliminated and replaced with more concise and easily understandable language. Similarly, synthesized speech is generated with a tone and speed that matches the user's emotions.
[0815] The generated audio data and adjusted text data are sent to the user's terminal and provided visually and audibly. Through this, the user can obtain an optimized learning experience. Furthermore, as the user provides feedback during the learning process, the server uses it to improve the entire system.
[0816] For example, when a user is reading a physics textbook, the server simplifies the text and adjusts the audio data based on acquired book data and real-time sentiment analysis, allowing the user to absorb the knowledge in a relaxed state.
[0817] An example of a prompt might be, "Analyze the user's facial expressions and voice to estimate their emotions and provide appropriate learning materials."
[0818] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0819] Step 1:
[0820] The server receives a book selection request sent from the user's terminal. The input is the title or identifier of the book selected by the user. The server sends a query to an external database to retrieve the book data. The retrieved book data is then output.
[0821] Step 2:
[0822] The server analyzes acquired book data using natural language processing techniques. The input is the text data of the books. Generative AI models (e.g., BERT, GPT) are used to analyze the complexity of the text and identify key topics. The output provides text feature information based on the analysis results.
[0823] Step 3:
[0824] The device captures the user's facial expressions and voice data. Inputs include real-time video and audio data. The device sends this data to the server. The server receives raw sensor data based on facial expressions and voice as output.
[0825] Step 4:
[0826] The server analyzes emotional data transmitted from the terminal. Facial and voice data are used as input. Standard facial recognition and voice analysis technologies are applied to identify the user's emotional state. Information regarding the user's emotional state is generated as output.
[0827] Step 5:
[0828] The server dynamically adjusts the text of book data based on the analyzed emotional state. Inputs are user emotional information and book feature information. A simplification algorithm is applied, eliminating technical jargon as needed and converting the text into concise language. The output is the adjusted text data.
[0829] Step 6:
[0830] The server performs speech synthesis based on simplified text. The input consists of new text data and parameters to reflect the user's emotions. This results in output speech data generated with a tone and speed appropriate to the emotional state.
[0831] Step 7:
[0832] The server sends the adjusted text data and generated audio data to the user's terminal. The input is this data, which can then be displayed and played back visually and audibly on the terminal. The output is content in a format usable by the user.
[0833] Step 8:
[0834] After completing the learning process, users input feedback into their device. This feedback includes their opinions and evaluations. The device sends this feedback to a server, and the output is data that is used to improve the system.
[0835] (Application Example 2)
[0836] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0837] Traditional learning systems have a problem in that they cannot flexibly respond to the individual emotional states of users, making it difficult to provide an appropriate learning environment. In particular, there was no mechanism to adjust the content and expression to match the user's emotions when they were feeling stressed or fatigued. As a result, users may experience decreased learning efficiency and a decline in motivation to learn.
[0838] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0839] In this invention, the server includes means for analyzing e-book data using natural language processing means to automatically identify the complexity and main topics of the text, means for analyzing the user's facial expressions and vocal state using emotion recognition means to identify emotions, and means for adjusting the difficulty level and expression of the text according to the emotional state. This makes it possible to adjust the learning content and expression to suit the user's emotional state.
[0840] "Natural language processing methods" are technologies that analyze e-book data and automatically identify the complexity and main topics of the text.
[0841] "User comprehension level" is an indicator that shows how well a user can understand a text, and it is used to adjust the learning content.
[0842] "Speech formation means" refers to a technology that generates speech information from simplified text.
[0843] A "user device" is a terminal used by a user, and is a device that enables the display and playback of audio information and simplified text.
[0844] "Emotion recognition means" refers to technology that identifies emotions by analyzing the user's facial expressions and vocal state.
[0845] "Emotional state" refers to the temporary emotional state a user is experiencing, and it is information that influences the difficulty level and expression adjustments of the learning content.
[0846] This invention realizes a system that provides an interactive learning environment that responds to the user's emotional state. First, the server retrieves e-book data from an external data store. This data is automatically analyzed using natural language processing to identify its complexity and main topics. Based on this, an algorithm is applied to simplify the text based on the user's level of understanding.
[0847] Next, the device collects data on the user's facial expressions and voice, and identifies emotions using emotion recognition tools. This process utilizes APIs such as Google Cloud Vision API and voice analysis API to obtain real-time emotion information from the user. Based on this obtained emotion information, the server adjusts the difficulty and expression of the simplified text, and generates voice information from the adjusted text using speech generation tools.
[0848] As a concrete example, if a user is reading a physics article on their smartphone and the device detects signs of fatigue, the server will respond by converting the text into simpler language, avoiding complex terminology, and generating voice information in a calmer tone. In this case, a speech synthesis API such as Amazon Polly would be used.
[0849] Finally, this adjusted text and generated audio information are sent to the user's device and played back as both text and audio. Through this, the user can receive a personalized learning experience, improving learning efficiency and motivation. The system is designed so that the generating AI model makes optimal adjustments based on a prompt message such as, "How should the content be simplified and the voice tone adjusted if the user is judged to be tired?"
[0850] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0851] Step 1:
[0852] The server retrieves e-book data from an external data store. This data is passed to a natural language processing system to analyze the complexity of the text and its main topics. The information obtained from this analysis (degree of complexity and main topics) becomes the output.
[0853] Step 2:
[0854] Based on the output of Step 1, the server applies an algorithm to simplify the text to match the user's level of understanding. The input consists of complexity information and the user's level of understanding, and the server generates a simplified text based on this. The resulting simplified document becomes the output.
[0855] Step 3:
[0856] The device collects the user's facial expressions and voice through its camera and microphone, and transmits this data to an emotion recognition system for analysis. The input is the user's facial expressions and voice data, and the emotion recognition result is output.
[0857] Step 4:
[0858] The server uses the sentiment recognition results obtained in step 3 to adjust the difficulty level and expression of the simplified text. The input is the simplified text and the sentiment recognition results, and the output is the adjusted text.
[0859] Step 5:
[0860] The server converts the adjusted text into speech data using speech synthesis means. The input is the adjusted text, and the speech data is output by speech synthesis.
[0861] Step 6:
[0862] The terminal receives audio data and edited text sent from the server, and displays and plays them for the user. The input is the data received from the server, and the output is the presentation of visual and auditory information to the user.
[0863] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0864] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0865] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0866] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0867] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0868] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0869] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0870] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0871] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0872] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0873] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0874] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0875] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0876] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0877] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0878] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0879] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0880] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0881] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0882] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0883] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0884] The following is further disclosed regarding the embodiments described above.
[0885] (Claim 1)
[0886] A means of analyzing e-book data using natural language processing to automatically identify the complexity of the text and its main topics,
[0887] A means of applying an algorithm to simplify text based on the user's reading comprehension level,
[0888] A means for generating speech data from simplified text using speech synthesis means,
[0889] A means for sending audio data and simplified text to a user terminal, enabling display and playback,
[0890] A means of collecting user feedback and using it to improve the system,
[0891] A system that includes this.
[0892] (Claim 2)
[0893] The system according to claim 1, comprising data acquisition means for acquiring ebook data, wherein the data acquisition means acquires data of a target book from an external database.
[0894] (Claim 3)
[0895] The system according to claim 1, further comprising means for improving simplification accuracy by analyzing collected feedback and optimizing the generation algorithm.
[0896] "Example 1"
[0897] (Claim 1)
[0898] The information processing device provides means for analyzing document data and automatically extracting the complexity of sentences and key topics,
[0899] A means of applying computational means to simplify a document based on the user's reading comprehension ability,
[0900] A means for generating audio information from a simplified document using a speech generation device,
[0901] A means for transmitting audio information and simplified documents to a user device, enabling display and playback,
[0902] A means of collecting feedback from users and using it to improve information processing methods,
[0903] A system that includes this.
[0904] (Claim 2)
[0905] The system according to claim 1, comprising information acquisition means for acquiring document data, wherein the information acquisition means acquires data of a target document from an external data storage device.
[0906] (Claim 3)
[0907] The system according to claim 1, further comprising means for improving simplification accuracy by analyzing collected opinions and optimizing the generation calculation means.
[0908] "Application Example 1"
[0909] (Claim 1)
[0910] A means of analyzing digital information data using natural language processing technology to automatically identify the difficulty level and important themes of a text,
[0911] A means of applying calculation methods to simplify the text according to the user's level of understanding,
[0912] A means of creating audio information from simplified text using speech generation technology,
[0913] A means for transmitting audio information and simplified text to a user terminal, enabling display and audio playback,
[0914] A means of aggregating user evaluation information and using it to improve system performance,
[0915] A means of providing personalized learning assistance to users by simplifying text and converting it to speech based on the digital information content selected by the user,
[0916] A system that includes this.
[0917] (Claim 2)
[0918] The system according to claim 1, comprising information acquisition means for acquiring digital information of a target book, wherein the information acquisition means acquires information to be analyzed from an external information base.
[0919] (Claim 3)
[0920] The system according to claim 1, further comprising means for improving simplification accuracy by analyzing aggregated evaluation information and optimizing the calculation algorithm.
[0921] "Example 2 of combining an emotion engine"
[0922] (Claim 1)
[0923] A means of analyzing document data using natural language processing and automatically identifying the complexity and main topics of the text,
[0924] A means of applying processing to simplify text based on the user's reading comprehension level,
[0925] To analyze the emotional state of users, we need a means to acquire and analyze emotional data,
[0926] A means of dynamically adjusting text and audio expression based on emotions,
[0927] A means for generating speech data from adjusted text using speech synthesis means,
[0928] A means for sending audio data and adjusted text to a user terminal, enabling display and playback,
[0929] A means of collecting feedback after training and using it to improve the accuracy of the system,
[0930] A system that includes this.
[0931] (Claim 2)
[0932] The system according to claim 1, comprising data acquisition means for obtaining document data from an external database, and evaluating the user's state through an interface which is emotion data acquisition means.
[0933] (Claim 3)
[0934] The system according to claim 1, further comprising means for improving the learning experience by analyzing collected feedback and sentiment data and optimizing speech generation and text simplification algorithms.
[0935] "Application example 2 when combining with an emotional engine"
[0936] (Claim 1)
[0937] A means of analyzing e-book data using natural language processing and automatically identifying the complexity of the text and its main topics,
[0938] A means of applying techniques to simplify text based on the user's level of understanding,
[0939] A means for generating audio information from simplified text using a speech formation means,
[0940] A means for transmitting audio information and simplified text to a user device, enabling display and playback,
[0941] A means of collecting user feedback and using it to improve the system,
[0942] An emotion recognition means analyzes the user's facial expressions and voice state to identify emotions,
[0943] A means of adjusting the difficulty and expression of a text according to the emotional state,
[0944] A system that includes this.
[0945] (Claim 2)
[0946] The system according to claim 1, comprising an information acquisition means for acquiring electronic book data, wherein the information acquisition means acquires information on a target book from an external data store.
[0947] (Claim 3)
[0948] The system according to claim 1, further comprising means for improving simplification accuracy by analyzing collected opinions and optimizing the generation method. [Explanation of symbols]
[0949] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of analyzing e-book data using natural language processing to automatically identify the complexity of the text and its main topics, A means of applying an algorithm to simplify text based on the user's reading comprehension level, A means for generating speech data from simplified text using speech synthesis means, A means for sending audio data and simplified text to a user terminal, enabling display and playback, A means of collecting user feedback and using it to improve the system, A system that includes this.
2. The system according to claim 1, comprising data acquisition means for acquiring ebook data, wherein the data acquisition means acquires data of a target book from an external database.
3. The system according to claim 1, further comprising means for improving simplification accuracy by analyzing collected feedback and optimizing the generation algorithm.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A