System
The system improves reading immersion by using gaze tracking and context analysis to select appropriate background music and provide on-the-spot explanations, addressing the issues of mismatched music and complex content understanding.
Patent Information
- Application Number
- JP2024123853
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Existing reading experiences are impaired by mismatched background music and difficulties in understanding complex words or contexts, leading to a lack of immersion and flow in reading.
A system that uses a front-facing camera and gaze tracking to analyze the reader's gaze and content, selecting appropriate background music and providing explanations for difficult words or contexts using a bone conduction speaker.
Enhances the reading experience by providing contextually relevant background music and immediate explanations, allowing readers to immerse themselves deeper in the book.
Smart Images

Figure 2026022336000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] "There is a problem that it is not possible to select appropriate background music while reading, and the reading experience is impaired because the music does not match the content of the book. There is also the problem that when reading, when you come across difficult words or context, it takes time to understand, and the flow of reading stagnates. The problem that this invention aims to solve is to solve these problems and provide a reading experience that allows readers to immerse themselves more deeply in the world of the book." [Means for solving the problem]
[0005] The present invention includes a means for detecting the reader's gaze direction using a camera installed in front of the device and a gaze measurement means for collecting the reader's gaze information in real time. It also includes an artificial intelligence means for analyzing the content being read based on the collected gaze information and data from the front camera, and a background music selection means for selecting appropriate background music based on the analyzed content. The system also includes a bone conduction speaker means for playing the selected background music, and an explanation generation means for generating and providing explanations of difficult words and contexts. This allows the reader to enjoy appropriate background music for the reading situation while being supported in understanding difficult parts, which is expected to improve the reading experience.
[0006] A "camera" is a device that captures the objects or text that a reader is looking at as image data.
[0007] An "eye gaze measurement device" is a system that detects the reader's gaze direction and viewpoint movements in real time and collects them as data.
[0008] "Artificial intelligence means" refers to software or a system that has the ability to analyze collected gaze information and camera data and determine what the reader is reading and their level of understanding.
[0009] A "background music selection means" is an algorithm or system that automatically selects appropriate background music based on the content being read.
[0010] A "bone conduction speaker" is a speaker device that transmits vibrations through the bones to the inner ear, allowing the user to hear sound.
[0011] An "explanation generator" is a system that generates meanings and summaries of difficult words and contexts and provides them to readers.
[0012] A "reading aid" is an electronic device designed to enhance the reader's reading experience, integrating the various means described above.
[0013] "Gaze information" is data that indicates which part the reader is looking at. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] The present invention provides a glasses-type reading assistance device for improving the reading experience. The device includes a front camera, a gaze measurement unit, an artificial intelligence unit, a background music selection unit, a bone conduction speaker, and a description generation unit. The program processing of this system is explained below in natural language.
[0036] 1. Setup Phase
[0037] The user puts on the "AI reading glasses." The device automatically starts up and activates the eye tracker to follow the user's gaze and the front camera to capture the text being read. The device is then ready for reading.
[0038] 2. Real-time understanding of reading status
[0039] The device captures the page of the book the user is looking at with the front camera and sends the information to the server. At the same time, the device continues to acquire gaze information using the eye tracker, and this information is continuously sent to the server.
[0040] 3. Analysis of reading content
[0041] The server receives the text and gaze data and analyzes it. This analysis identifies which part the user is reading and understands what they are reading. Based on this information, the server can select appropriate background music or detect difficult passages.
[0042] 4. Select and play appropriate background music
[0043] The server selects appropriate background music based on the reading content analyzed. The selected background music data is sent to the device, which then provides the background music to the user using a bone conduction speaker. This allows music to be played according to the situation, enhancing the reading atmosphere.
[0044] 5. Assistance with difficult-to-read passages
[0045] If the server determines that a particular word or context is difficult for the user to understand, it generates a meaning or summary of that word or context, sends this information to the device, and provides it to the user audibly through a bone conduction speaker. If necessary, it may also be displayed as visual feedback on the feedback panel.
[0046] Specific operation example
[0047] Reading situations
[0048] A user is reading a historical novel and is about to enter a battle scene. The device captures the scene through the front camera and acquires gaze information.
[0049] The server receives the captured text and gaze information and analyzes it to determine that it is a battle scene.
[0050] The server selects epic battle music that heightens the tension and sends it to the device.
[0051] The device provides the music to the user through a bone conduction speaker.
[0052] If the difficult word "Yamato Taro Yoshiie" appears in the middle of a battle scene,
[0053] The server determines that this word is difficult for the user to understand and generates a meaning and background such as "Hachiman Taro Yoshiie was a military commander in the Heian period, and was later enshrined as a god at Hachiman Shrine."
[0054] The device provides this information to the user through a bone conduction speaker.
[0055] This allows users to immerse themselves even more deeply in the world of the book through music and explanations.
[0056] As can be seen, the present invention enhances the reading experience and allows readers to become immersed in the world of the book.
[0057] The processing flow will be explained below.
[0058] Step 1:
[0059] The user puts on the "AI reading glasses." The device automatically starts up and activates the eye tracker to follow the user's gaze and the front camera to capture the text being read.
[0060] Step 2:
[0061] The device uses the front camera to capture the page of the book the user is looking at, formats the captured image data, and saves it in temporary storage.
[0062] Step 3:
[0063] The device uses an eye tracker to acquire user gaze information and temporarily stores the gaze data.
[0064] Step 4:
[0065] The device sends the captured image data and gaze data as a set to the server in real time.
[0066] Step 5:
[0067] The server converts the received image data into text data through an OCR (optical character recognition) process. Based on the converted text data and gaze data, the server performs natural language analysis to identify what the user is reading and where they are reading from.
[0068] Step 6:
[0069] Based on the analyzed text data, the server selects the appropriate background music for the scene. For example, if it's a battle scene, it will select epic battle music.
[0070] Step 7:
[0071] The server sends the selected background music data to the device, which then plays it through the bone conduction speaker, allowing users to enjoy music suitable for reading.
[0072] Step 8:
[0073] The server uses gaze data and text data to detect words and contexts that are difficult to understand, such as difficult historical or academic terms.
[0074] Step 9:
[0075] The server generates meaning and context for each difficult word and context that is detected, which is then converted into audio data.
[0076] Step 10:
[0077] The server sends the generated audio data to the device, which then plays it back through a bone conduction speaker and optionally displays visual information on a feedback panel.
[0078] Through the above steps, the present invention improves the reading experience and provides an environment where users can immerse themselves in the world of books.
[0079] Example 1
[0080] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0081] The modern reading experience remains static, making it difficult for readers to deeply understand the meaning of text or customize their reading environment. Furthermore, there is a lack of ways to quickly understand the meaning of difficult words or contexts when encountering them. Furthermore, there is no function to select appropriate background music, making it difficult to create an environment conducive to immersion in reading.
[0082] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0083] In this invention, the server includes means for detecting the user's gaze direction using a camera installed in front, gaze tracking means for collecting the user's gaze information in real time, machine learning means for analyzing the content being read based on the collected gaze information and data from the camera, background sound selection means for selecting appropriate background sound based on the analyzed content, bone conduction sound means for playing the selected background sound, and explanation generation means for generating and providing explanations of difficult words and contexts as needed, thereby improving the reading experience, enabling the reader to deeply understand the content of the book, and providing an appropriate acoustic environment.
[0084] The "photography device" is a device installed in front of the user to detect the direction of the user's line of sight.
[0085] "Eye tracking means" is a technology for collecting user eye gaze information in real time.
[0086] "Machine learning means" is a technology that analyzes the content of reading based on collected gaze information and data from a camera.
[0087] The "background sound selection means" is a technique for selecting appropriate background sound based on the analyzed content.
[0088] "Bone conduction sound means" is a technology for reproducing selected background sounds.
[0089] "Explanation generator" is a technology for generating and providing explanations of difficult words and contexts as needed.
[0090] The present invention relates to a glasses-type reading assistance device for improving reading experience, which includes a front-mounted camera, an eye tracking unit, a machine learning unit, a background sound selection unit, a bone conduction sound unit, and an explanation generation unit.
[0091] When a user puts on the eyeglass-type reading assistance device, the device automatically starts up and activates the eye tracker (eye tracking means) to track the user's gaze and the front camera (photography device) to capture the text being read, making the device ready for reading.
[0092] Next, the device captures the page of the book the user is looking at with the front camera and sends the information to the server. At the same time, the device continues to acquire gaze information using the eye tracker, which is then continuously sent to the server.
[0093] The server analyzes the transmitted gaze data and captured text images. It uses machine learning methods to identify which part the user is reading and understand what is being read. Based on this information, the server can select appropriate background audio or detect difficult words or contexts.
[0094] The server then selects appropriate background sounds based on the analysis results. The selected background sound data is sent to the device, which then provides the sound to the user using bone conduction audio, thereby playing music appropriate to the situation and enhancing the reading atmosphere.
[0095] Furthermore, if the server determines that a particular word or sentence is difficult for the user to understand as a result of the analysis, it generates a meaning or summary of the word or sentence. This information is sent to the terminal, which then provides it to the user audibly through bone conduction acoustic means. If necessary, the terminal may also display the information on a feedback panel as visual feedback.
[0096] Specific examples
[0097] For example, suppose a user is reading a historical novel. When the user comes to a battle scene in the book, the device captures the scene through the front camera and acquires gaze information. The server receives the captured text and gaze information and analyzes that it is a battle scene. As a result, the server selects epic battle sounds to heighten the tension and transmits them to the device. The device then provides this sound to the user through bone conduction audio.
[0098] Also, if the user comes across a difficult word like "Yoshiie Hachiman" while reading, the server will determine that the word is difficult for the user to understand and generate a meaning and background such as "Yoshiie Hachiman was a military commander in the Heian period, and was later enshrined as a god at Hachiman Shrine." The terminal will provide this information to the user through bone conduction acoustic means.
[0099] Prompt Sentence Examples
[0100] "Please select background music that matches the battle scenes in a historical novel and generate commentary for Yawata Taro Yoshiie."
[0101] As described above, the present invention improves the reading experience, allows the reader to understand the contents of the book more deeply, and provides an appropriate acoustic environment.
[0102] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0103] Step 1:
[0104] The user puts on the "glasses-type reading assistance device." The device automatically starts up, and the eye tracker (a means of tracking gaze) and the front camera (a photographing device) for capturing the text being read are enabled. This allows the device to obtain gaze information in real time, preparing the user's reading environment.
[0105] Step 2:
[0106] The device captures the page of the book the user is looking at through the front camera. At the same time, the eye tracker collects the user's gaze information. As input, the device receives image data and gaze data from the front camera, and sends these data to the server in real time. As output, the captured image data and gaze data are sent to the server.
[0107] Step 3:
[0108] The server analyzes the received gaze data and captured image. It uses gaze data and captured image as input. It uses machine learning algorithms to identify which part the user is reading and performs text analysis. The output is the specific information about the reading part and the analysis results.
[0109] Step 4:
[0110] The server selects appropriate background audio based on the analysis results. It uses the specific information about the reading section and the analysis results as input. The server uses a background audio selection algorithm to select music that matches the analyzed content. The selected background audio data is generated as output.
[0111] Step 5:
[0112] The server sends the selected background sound data to the terminal. The background sound data is used as input. The terminal receives this sound data and provides music to the user using bone conduction sound means. As output, appropriate background music is played for the user.
[0113] Step 6:
[0114] The server identifies difficult words and sentences and generates their meaning and context. It uses gaze data and text analysis results as input. It uses an explanation generation algorithm to create explanations for difficult words and contexts. The output is the generated explanation data.
[0115] Step 7:
[0116] The server sends the generated explanation data to the terminal, which uses the explanation data as input. The terminal receives the explanation data and provides audio explanations to the user through bone conduction acoustic means, and also provides visual feedback if necessary. As output, the user is provided with audio and visual explanations of difficult words and contexts.
[0117] Through the above processing steps, the present invention can improve the reading experience, enable users to understand the contents of the book more deeply, and provide a suitable acoustic environment.
[0118] (Application example 1)
[0119] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0120] Traditional reading experiences have the drawback of making it difficult for readers to gain a deep understanding of the content or experience a sense of realism. In particular, when encountering difficult words or context, there are limited ways to receive on-the-spot explanations, which frequently interrupts the flow of reading. There is also a demand for incorporating background music and sound effects into reading to improve the quality of the reading experience.
[0121] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0122] In this invention, the server includes: means for detecting the reader's gaze direction using a camera installed in front; gaze measurement means for collecting the reader's gaze information in real time; artificial intelligence means for analyzing the content being read based on the collected gaze information and data from the front camera; background music selection means for selecting appropriate background music based on the analyzed content; a sound wave conduction audio device for playing the selected background music; explanation generation means for generating and providing explanations of difficult words and context as needed; data communication means for transmitting the text and gaze information being read to the data server, the server analyzing the transmitted data, and providing the user with the analysis results; and audio control means for transmitting the background music selected based on the analysis results to the smart device and for the smart device to play the background music. This significantly improves the quality of the reading experience, allowing the reader to deeply understand the content and experience a sense of realism.
[0123] 1. "Camera" refers to a device that is installed in front of the user to detect the direction of their gaze and capture images in real time.
[0124] 2. "Gaze measurement means" refers to technology or equipment used to collect readers' gaze information in real time and analyze that information.
[0125] 3. "Artificial intelligence means" refers to algorithms and processes used to analyze the content of a reading based on collected gaze information and camera data.
[0126] 4. "Background music selection means" refers to a method or system for selecting appropriate background music based on the analyzed reading content.
[0127] 5. "Sound wave conduction sound device" means a bone conduction speaker or similar sound reproduction device that provides a user with selected background music.
[0128] 6. "Explanation generation means" refers to technology that generates explanations for difficult words or contexts and provides them to users in audio or visual form.
[0129] 7. "Data communication means" refers to the network technology and protocols used to transmit text and eye gaze information during reading to a data server and receive analysis results.
[0130] 8. "Audio Control Means" means the control technology or algorithm for playing background music selected based on the analysis results on a smart device.
[0131] 9. "Data server" refers to a computer system installed on the cloud or elsewhere that analyzes transmitted data and provides the results.
[0132] To implement this invention, the following system configuration and its operating procedures will be specifically described. The system includes a camera, a gaze measurement means, an artificial intelligence means, a background music selection means, a sound-transmitting audio device, a description generation means, a data communication means, and an audio control means.
[0133] Hardware and Software Use
[0134] Camera: Uses the front-facing camera on your smartphone or tablet to capture where you're looking and the text you're reading.
[0135] Eye tracking method: Use software with eye tracking functionality (e.g., the GazeTracking library).
[0136] Artificial intelligence measures: Machine learning models are used to analyze collected gaze information and captured text data.
[0137] Background music selection method: Use a music library and music recommendation algorithms to select background music appropriate for the reading content.
[0138] Sound wave conduction sound device: Plays background music using bone conduction speakers or smartphone speakers.
[0139] Explanation generation: Use text generation models and text-to-speech technologies (e.g., Google Cloud Text-to-Speech API) to generate explanations for difficult words and contexts and convert them into audio.
[0140] Data communication means: An internet communication protocol (e.g., HTTP request) is used to send text data and gaze data to the cloud server and return the analysis results to the smart device.
[0141] Sound control means: Use the music playback library to play background music on the smart device based on the analysis results sent from the cloud server.
[0142] Specific examples
[0143] Reading situations
[0144] Assume a user is reading a historical novel on their smartphone. When they reach a battle scene in the book, the system works as follows:
[0145] 1. Eye Tracking and Text Capture:
[0146] The page the user is reading is captured through the smartphone camera.
[0147] At the same time, the gaze measurement means collects the user's gaze data in real time.
[0148] 2. Data analysis and background music selection:
[0149] The captured text data and gaze data are transmitted to a cloud server.
[0150] The cloud server uses machine learning models to analyze the data and identify that the user is reading a fight scene.
[0151] Based on the analysis results, epic battle music is selected and the music data is sent back to the smartphone.
[0152] 3. Play background music and explain difficult passages:
[0153] The smartphone plays battle music through a sonic wave transmission audio device.
[0154] If the difficult word "Yahata Taro Yoshiie" appears in the middle of a battle scene, the cloud server will determine that the word is difficult to understand and generate its meaning and background.
[0155] The smartphone plays the generated explanation as audio and provides it to the user.
[0156] Prompt Sentence Examples
[0157] A user is reading a historical novel and is currently approaching a battle scene. Choose epic battle music to match this scene and explain the meaning and background of the difficult passage "Yaman Taro Yoshiie."
[0158] This will significantly improve the quality of the reading experience, allowing readers to gain a deeper understanding of the content and experience a more immersive experience.
[0159] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0160] Step 1:
[0161] The user opens the smartphone, launches the app, and selects an e-book. The input is the e-book file selected by the user and the initial gaze information data. The output is a state where camera capture and gaze tracking are ready. The camera is placed in front and the eye tracking function is enabled.
[0162] Step 2:
[0163] The device uses a front camera to capture text data of the page the user is reading and collects gaze data in real time through an eye gaze measurement device. The input is the camera image and real-time gaze position data. The output is the captured text data and gaze data.
[0164] Step 3:
[0165] The captured text data and gaze data are sent from the terminal to the server. The input is the text data and gaze data sent from the terminal. The output is a notification that the text data and gaze data have been received. These data are sent to the cloud server using a data communication means.
[0166] Step 4:
[0167] The server analyzes the received text data and gaze data using artificial intelligence. The input is the text data and gaze data. The output is the analysis result of what the user is reading. For example, it may be determined that the user is reading a battle scene.
[0168] Step 5:
[0169] The server selects appropriate background music based on the analysis results. The input is the analyzed reading content (e.g., a battle scene). The output is a URL or file of the selected background music data. Using the background music selection means, epic music suitable for the battle scene is selected.
[0170] Step 6:
[0171] The server sends the selected background music data to the terminal. The input is the URL or file of the background music data. The output is the background music data sent to the terminal. The music data is sent to the smartphone using a data communication means.
[0172] Step 7:
[0173] The background music data received by the terminal is played on an acoustic wave conduction acoustic device. The input is the received background music data. The output is the background music being listened to by the user. Using the acoustic control means, the background music is played on the smartphone's bone conduction speaker.
[0174] Step 8:
[0175] The server identifies difficult-to-read passages and generates their meaning and background using an explanation generation means. The input is the analyzed text data and information on the identified difficult-to-read passages. The output is the generated explanatory text. For example, the text generated for "Hachiman Taro Yoshiie" is "He was a military commander in the Heian period, and later enshrined as a god at Hachiman Shrine."
[0176] Step 9:
[0177] The server transmits the explanation generated by the explanation generation means to the terminal. The input is the generated explanation text. The output is the explanation audio data transmitted to the terminal. The explanation data is transmitted to the smartphone using the data communication means.
[0178] Step 10:
[0179] The terminal plays the received explanatory data on an acoustic wave conduction audio device. The input is the received explanatory audio data. The output is the explanatory audio that the user is listening to. Using the audio control means, the explanation is played through a bone conduction speaker.
[0180] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0181] This invention is a glasses-type reading assistance device that combines a front camera, gaze tracking means, artificial intelligence means, background music selection means, bone conduction speakers, explanation generation means, and an emotion engine to improve the reading experience. This device analyzes the content being read based on the user's gaze and emotions, plays background music appropriate to the situation, and provides support for difficult words and context. The program processing of this system is explained below in natural language.
[0182] 1. Setup Phase
[0183] The user puts on the AI reading glasses. The device automatically activates the eye tracker to track their gaze and the front camera to capture the text they are reading. It also activates an emotion engine that analyzes the user's facial expressions and tone of voice.
[0184] 2. Real-time understanding of reading status
[0185] The device captures the page of the book the user is looking at with the front camera and sends the information to the server. At the same time, the device continues to acquire gaze information using the eye tracker. The emotion engine also analyzes the user's facial expressions and tone of voice to acquire emotional data. This information is continuously sent to the server.
[0186] 3. Analysis of reading content
[0187] The server receives the text, gaze data, and emotion data and analyzes them. The server recognizes the text data using OCR (optical character recognition) technology and combines it with the gaze data to identify the specific part the user is reading. It also analyzes the emotion data to determine the user's current emotional state.
[0188] 4. Select and play appropriate background music
[0189] The server selects appropriate background music based on the reading content and the user's emotional data analyzed. For example, if the user is feeling tense while reading, it selects relaxing music, and if the user is enjoying themselves, it selects music that enhances the mood. The selected background music data is sent to the device, which then uses a bone conduction speaker to provide the background music to the user.
[0190] 5. Assistance with difficult-to-read passages
[0191] If the server detects words or contexts that are difficult to read from the gaze data and text data, it generates a meaning and summary of the words or contexts. The generated information is converted into audio data and sent from the server to the device. The device then provides the audio data to the user through a bone conduction speaker. It may also display visual information on a feedback panel as needed.
[0192] Specific operation example
[0193] Reading situations
[0194] A user is reading a mystery novel. The climax of the story is approaching, and the user's heart rate is rising. The device captures the scene through the front camera and acquires gaze and emotion information.
[0195] The server receives the captured text, gaze information, and emotion data and analyzes it to determine whether it is a climax scene.
[0196] The server selects relaxing music to relieve tension based on the user's emotional state and sends it to the terminal.
[0197] The device provides the music to the user through a bone conduction speaker.
[0198] When the difficult name "Albaiohiko" appears in the climax scene,
[0199] The server determines that this name is difficult for the user to understand and generates a meaning and background, such as "Albaiohiko is a mysterious detective and a character with a hidden past."
[0200] The device provides this information to the user through a bone conduction speaker.
[0201] This allows users to receive appropriate music and explanations while reading, allowing them to become more immersed in the world of the book.
[0202] As described above, the present invention improves the reading experience and provides an environment in which users can immerse themselves in the world of books.
[0203] The processing flow will be explained below.
[0204] Step 1:
[0205] The user puts on the "AI reading glasses" and turns on the device, which activates the eye tracker, front-facing camera, and emotion engine.
[0206] Step 2:
[0207] The device uses the front camera to periodically capture the page of the book the user is looking at, and the captured image data is temporarily stored.
[0208] Step 3:
[0209] The device uses an eye tracker to acquire the user's gaze information in real time and temporarily stores this gaze data.
[0210] Step 4:
[0211] The device uses an emotion engine to analyze the user's facial expressions and tone of voice to determine their current emotional state, and emotion data is also stored.
[0212] Step 5:
[0213] The terminal transmits the captured image data, gaze data, and emotion data as a set to the server in real time.
[0214] Step 6:
[0215] The server converts the received image data into text data using OCR (optical character recognition) technology, and then analyzes the text data by combining it with gaze data and emotion data.
[0216] Step 7:
[0217] Based on the analysis, the server identifies the specific part the user is reading and makes a comprehensive assessment of the user's emotional state.
[0218] Step 8:
[0219] The server selects appropriate background music based on the analysis results and emotional data. For example, it selects relaxing music for tense scenes and music that enhances the atmosphere for happy scenes.
[0220] Step 9:
[0221] The server sends the selected background music data to the device, which then uses a bone conduction speaker to provide the music to the user.
[0222] Step 10:
[0223] The server detects difficult words and contexts from gaze data and text data, and generates meaning and background information for the detected content.
[0224] Step 11:
[0225] The server converts the generated explanation data into audio data and sends it to the terminal, which then provides the audio data to the user through a bone conduction speaker. The terminal also displays visual information on a feedback panel as needed.
[0226] As a specific example of how it works, when a user is reading a mystery novel and reaches a tense climax, the emotion engine detects the user's tension.
[0227] The server selects relaxing background music and sends it to the device.
[0228] The device plays the background music through a bone conduction speaker, helping the user relax.
[0229] When the difficult name "Albaiohiko" appears, the server generates background information about the name and sends it to the terminal as voice data.
[0230] The device plays audio to help the user understand.
[0231] In this way, the present invention takes into account the user's emotions while providing appropriate music and information support to enhance the reading experience.
[0232] Example 2
[0233] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0234] While various technologies have been proposed to improve the reading experience, systems that provide real-time support based on changes in the reader's gaze and emotions are not yet fully developed. In particular, technologies that provide appropriate background music while reading and clearly explain difficult words and contexts are in need of further development. This creates a demand for systems that can enhance immersion and comprehension while reading.
[0235] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0236] In this invention, the server includes means for detecting the reader's gaze direction using an image sensor installed in front, a gaze measurement device for collecting the reader's gaze information in real time, artificial intelligence means for analyzing the content being read based on the collected gaze information and data from the front image sensor, background music selection means for selecting appropriate background music based on the analyzed content and the reader's emotional data, a bone conduction audio device for playing the selected background music, explanation generation means for generating and providing explanations of difficult terms and context as necessary, and voice conversion means for converting the analyzed data into voice data and providing it to the reader. This makes it possible to provide dynamic support in real time according to the reader's gaze and emotions, significantly improving the quality of the reading experience.
[0237] A "forward-mounted image sensor" is an image capture device mounted in front of the reader to detect the reader's line of sight.
[0238] The "means for detecting the reader's line of sight" is a mechanism that uses an image sensor installed in front to identify where the reader is looking.
[0239] An "eye gaze measurement device that collects reader gaze information in real time" is a device that instantly tracks the reader's eye movements and collects that data.
[0240] "Artificial intelligence means" refers to technology that analyzes collected gaze information and data from the forward image sensor to understand the content and behavior of the reader while reading.
[0241] The "background music selection means" is a mechanism for selecting optimal background music based on the analyzed content and emotional data.
[0242] A "bone conduction audio device" is an audio playback device that allows selected background music to be heard directly through the bones.
[0243] "Explanation generation means" is a technology that generates the meaning and summary of text to provide easy-to-understand explanations of difficult terms and context.
[0244] "Speech conversion means" refers to a technology that converts the generated commentary and information into audio data and provides it to the reader.
[0245] The present invention is a glasses-type device system designed to enhance the reading experience. The device combines a front-mounted image sensor, an eye tracking device, an artificial intelligence means, a background music selection means, a bone conduction sound device, a commentary generation means, and a voice conversion means to provide optimal reading support to users. Detailed embodiments of the system are described below.
[0246] When a user puts on the AI reading glasses, the device automatically starts up. The device includes an eye tracker to track the user's gaze and a forward-facing image sensor to capture the text being read. An emotion engine then starts up, analyzing the user's facial expressions and tone of voice.
[0247] The device captures the page of the book the user is looking at with a front image sensor and sends the information to the server. At the same time, it acquires gaze information using an eye tracker and collects emotion data with an emotion engine. This information is continuously sent to the server.
[0248] The server converts the image data into text using OCR (optical character recognition) technology, such as software like Tesseract OCR. The acquired text data is combined with gaze data to identify the user's current reading position. At the same time, a sentiment analysis algorithm is used to analyze the emotion data and determine the user's current emotional state.
[0249] The server selects appropriate background music based on the analyzed content. For example, if you are feeling tense while reading, it will select relaxing music, and if you are enjoying yourself, it will select music that will enhance the atmosphere. The background music data is sent to the terminal and provided to the user via a bone conduction audio device.
[0250] Additionally, if the server detects words or contexts that are difficult to understand from the gaze data and text data, it generates a meaning or summary of those words or contexts. This content is generated using a generative AI model (such as the GPT series). The generated information is converted into audio data and sent from the server to the device. The device then provides the audio data to the user through a bone conduction audio device. If necessary, the information is also displayed on a visual feedback panel.
[0251] As a concrete example, consider a case where a user is reading a mystery fiction book. As the climax approaches, the user's heart rate increases. The device captures this scene through the forward-facing image sensor and acquires gaze information and emotional information.
[0252] The server receives the captured text, gaze information, and emotion data and analyzes it to determine whether it is a climax scene.
[0253] The server selects relaxing music to relieve tension based on the user's emotional state and sends it to the terminal.
[0254] The terminal provides the music to the user through a bone conduction audio device.
[0255] If a mysterious character called "Detective Albion" appears in the climax scene,
[0256] The server determines that this character is difficult for the user to understand, and uses a generative AI model to generate meaning and background, such as "Detective Albaio is a mysterious detective and a character with a hidden past."
[0257] The terminal provides this information to the user through a bone conduction audio device.
[0258] This system allows users to receive appropriate music and explanations while reading, allowing them to become more immersed in the world of the book. In addition, by linking it with a generative AI model, it is possible to generate prompts to aid the reader in comprehension.
[0259] Prompt Sentence Examples
[0260] Please provide a brief background on the character "Detective Albio."
[0261] Choose relaxing music to listen to while you read.
[0262] The above is an embodiment of the present invention, which allows users to have a more comfortable and immersive reading experience.
[0263] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0264] Step 1:
[0265] The user puts on the "AI Reading Glasses." The device automatically starts up, activating the eye tracker to track the user's gaze and the forward-facing image sensor to capture the text being read. The emotion engine also starts up.
[0266] Specifically, the device starts up the image sensor and eye tracker, and goes through an initialization process to prepare for acquiring gaze and video data.
[0267] Input: The user puts on the device.
[0268] Output: Eye tracking and image acquisition ready.
[0269] Step 2:
[0270] The device uses a forward-facing image sensor to capture the page of the book the user is reading, an eye tracker to collect gaze information in real time, and an emotion engine to extract emotional data from facial expressions and tone of voice.
[0271] Specifically, the device continuously captures images of the page using a front image sensor, acquires user gaze data using an eye tracker, and performs facial expression and voice analysis using an emotion engine.
[0272] Input: User's gaze information, forward image, emotion data.
[0273] Output: Captured image data, gaze data, emotion data.
[0274] Step 3:
[0275] The device continuously transmits captured image data, gaze data, and emotion data to the server.
[0276] Specifically, the terminal compresses the collected data at regular intervals and transmits it to the server via packet communication. Since the data is transmitted in real time, network bandwidth and communication efficiency are taken into consideration.
[0277] Input: Captured image data, gaze data, emotion data.
[0278] Output: Various data sent to the server.
[0279] Step 4:
[0280] The server converts the received image data into text using OCR technology (e.g., Tesseract OCR), combines it with gaze data to identify the part the user is reading, and analyzes emotion data to determine the user's emotional state.
[0281] Specifically, the server converts image data into text using an OCR engine, maps it with gaze data, and then uses an emotion analysis algorithm to identify the user's emotional state.
[0282] Input: Transmitted image data, gaze data, emotion data.
[0283] Output: Transformed text data, identified reading passages, and the user's emotional state.
[0284] Step 5:
[0285] The server selects appropriate background music based on the analysis results. For example, if the user is feeling tense, it selects relaxing music, and if the user is having fun, it selects music that will liven up the atmosphere. The selected background music data is sent to the device.
[0286] Specifically, the server selects the most suitable music from multiple music databases, compresses the music data, and sends it to the terminal.
[0287] Input: Parsed reading content, user's emotional state.
[0288] Output: Selected background music data, sent to the device.
[0289] Step 6:
[0290] The background music data received by the terminal is played through a bone conduction audio device.
[0291] Specifically, the device stores background music data in a buffer and plays it in real time through a bone conduction audio device.
[0292] Input: BGM data sent from the server.
[0293] Output: Providing background music to the user.
[0294] Step 7:
[0295] The server uses gaze data and text data to identify words and contexts that the user has difficulty comprehending, and generates meanings and summaries of those words and contexts. The generated information is converted into audio data and sent to the device.
[0296] Specifically, the server analyzes gaze data and text data, generates meaning and summaries using a generative AI model (e.g., GPT-3), and converts this text data into audio data.
[0297] Input: gaze data, text data.
[0298] Output: Generated description data and audio data.
[0299] Step 8:
[0300] The audio data received by the terminal is provided to the user through a bone conduction audio device.
[0301] Specifically, the device decodes the audio data, plays the audio through a bone conduction audio device, and optionally displays visual information on the feedback panel.
[0302] Input: Audio data sent from the server.
[0303] Output: Providing audio information to the user.
[0304] In this way, through the specific processing performed at each step, the user can receive appropriate reading assistance in real time.
[0305] (Application example 2)
[0306] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0307] Conventional reading assistance devices and electronic payment systems are unable to reflect the user's emotional state or gaze data in real time, limiting their ability to provide diverse information and improve the user experience. Furthermore, when shopping online, users often feel stressed when choosing the most suitable product and payment method, which can reduce their motivation to purchase. There is a need to solve these issues and provide users with a richer experience and more efficient shopping support.
[0308] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0309] In this invention, the server includes: means for detecting the user's gaze direction using a camera installed in front; gaze measurement means for collecting user gaze information in real time; artificial intelligence means for analyzing the content of reading and online activities based on the collected gaze information and data from the front camera; background music selection means for selecting appropriate background music based on the analyzed content; bone conduction speaker means for playing the selected background music; explanation generation means for generating and providing explanations of difficult words and contexts as necessary; product recommendation means for capturing the product page the user is viewing while shopping online and recommending appropriate products and services based on that information and emotional data; and payment assistance means for suggesting the optimal payment method based on the analyzed emotional data and supporting purchasing behavior. This enables optimal music selection and product recommendation based on emotional analysis.
[0310] A "front camera" is a photographic device installed on the front of a device to capture the user's line of sight and what they are viewing.
[0311] "Gaze measurement means" refers to technology or equipment that collects a user's gaze information in real time and analyzes its direction and focus.
[0312] "Artificial intelligence means" refers to technologies and systems that analyze collected gaze information and data from the front camera to understand the content of reading and online activities.
[0313] The "background music selection means" refers to a technology or system for selecting optimal background music based on the analyzed content.
[0314] "Bone conduction speaker means" means an audio device that uses bone conduction technology to transmit selected background music to a user.
[0315] "Explanation generation means" refers to technology or systems that generate explanations of difficult words or contexts as needed and provide them to users.
[0316] A "product recommendation method" is a technology or system that captures the product page a user is viewing while shopping online and recommends appropriate products or services based on that information and emotional data.
[0317] "Payment assistance methods" are technologies and systems that suggest the most suitable payment method to users based on analyzed emotional data and assist them in their purchasing behavior.
[0318] The present invention is implemented as a system for improving the reading experience and online shopping experience by a device that combines a forward-facing camera, a gaze measurement means, an artificial intelligence means, a background music selection means, a bone conduction speaker means, a description generation means, a product recommendation means, and a payment assistance means.
[0319] 1. Hardware configuration:
[0320] The front-facing camera is located at the front of the device to capture the user's line of sight and what they are viewing.
[0321] The gaze measurement means is used to collect user gaze information in real time.
[0322] The bone conduction speaker is used to convey selected background music and explanatory information to the user.
[0323] 2. Software configuration:
[0324] The artificial intelligence means analyzes the collected gaze information and data from the front camera to understand the content of your reading and online activities.
[0325] The background music selection means selects appropriate background music based on the analyzed content.
[0326] The explanation generating means generates explanations of difficult words and contexts as needed and provides them to the user.
[0327] The product recommendation means captures the product page that the user is viewing while shopping online and recommends appropriate products and services based on that information and emotional data.
[0328] The payment assistance tool suggests the most suitable payment method to the user based on the analyzed emotional data, and supports purchasing behavior.
[0329] 3. Data processing and calculation:
[0330] The video data acquired from the front camera is sent to a server and converted into text data using OCR technology.
[0331] The gaze data and emotion analysis data acquired by the gaze measurement means are analyzed by the artificial intelligence means to identify the specific part the user is reading.
[0332] The background music selection means selects the most suitable music based on the emotion and theme of the analyzed text and plays it through a bone conduction speaker.
[0333] The product recommendation and payment assistance methods capture the product page the user is viewing while shopping online, and suggest the most suitable products and payment methods based on that information and emotional data.
[0334] 4. Specific examples:
[0335] For example, if a user is reading a mystery novel, emotion analysis can tell if the user is feeling tense. The system can detect this and play relaxing music through a bone conduction speaker. It can also generate and provide audio explanations of difficult words and context that appear in the climax scene.
[0336] 5. Example of a generative AI model and prompt:
[0337] Example prompt sentence:
[0338] Capture the product page the user is viewing with a camera, collect gaze and emotion data, and send that data to your server to recommend the best products and payment methods.
[0339] This allows users to receive optimal music and descriptions, product recommendations, and payment assistance while reading or shopping online, greatly improving their experience.
[0340] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0341] Step 1:
[0342] The user puts on the device ("Smart Payment Assist Glasses") and the system starts automatically. The input is the user's wearing information. The device activates the camera and gaze measurement means and begins to capture the user's gaze information and facial expressions in real time. The output obtained is gaze data and facial expression data captured in real time.
[0343] Step 2:
[0344] The device uses the front camera to capture an image of the product page the user is viewing. The input is the product page the user is viewing. The image data captured by the camera is sent to the server. The output is the image data sent to the server.
[0345] Step 3:
[0346] The server analyzes the received image data using OCR (Optical Character Recognition) technology to extract text data related to the product. The input is image data captured by a camera. Through OCR analysis, text data is extracted to obtain the specific product information the user is viewing. The output is the extracted text data.
[0347] Step 4:
[0348] The device uses gaze measurement means to collect the user's gaze information in real time. The input is the user's eye movement and focus. The gaze data is continuously sent to a server and used to identify the specific location the user is looking at. The output is the identified gaze information.
[0349] Step 5:
[0350] The server uses an emotion engine to analyze the user's facial expressions and tone of voice to obtain current emotion data. The inputs are the user's facial expressions and tone of voice. The emotion data is analyzed along with the gaze data to help identify the user's current emotional state. The output is the identified emotion data.
[0351] Step 6:
[0352] The server comprehensively analyzes gaze data, text data, and emotion data to determine the products in which the user is interested and the degree of their willingness to purchase them. The inputs are gaze data, text data, and emotion data. The analysis results in the recommendation of appropriate products and services. The output is information about recommended products and services.
[0353] Step 7:
[0354] Based on the analyzed content, the server selects appropriate background music using a background music selection means. The inputs are the user's emotional state and the text content being viewed. The selected music is sent to the terminal and played to the user through a bone conduction speaker. The resulting output is the selected background music.
[0355] Step 8:
[0356] The server generates explanations of difficult words and contexts for the user and converts them into audio data. The input is the text content the user is viewing and the identified difficult parts. The explanation generation means generates a summary and sends it to the terminal as audio data. The output obtained is explanatory information as audio data.
[0357] Step 9:
[0358] The server proposes the optimal payment method based on the analyzed emotional data and purchasing intention information and notifies the user. The input is the user's emotional state and recommended product information. The payment assistance means selects the appropriate payment method and provides assistance to the user. The output obtained is the proposed payment method.
[0359] This allows users to receive the best music and explanations, product recommendations, and payment assistance while reading or shopping online.
[0360] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0361] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0362] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0363] [Second embodiment]
[0364] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0365] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0366] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0367] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0368] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0369] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0370] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0371] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0372] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0373] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0374] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0375] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0376] The present invention provides a glasses-type reading assistance device for improving the reading experience. The device includes a front camera, a gaze measurement unit, an artificial intelligence unit, a background music selection unit, a bone conduction speaker, and a description generation unit. The program processing of this system is explained below in natural language.
[0377] 1. Setup Phase
[0378] The user puts on the "AI reading glasses." The device automatically starts up and activates the eye tracker to follow the user's gaze and the front camera to capture the text being read. The device is then ready for reading.
[0379] 2. Real-time understanding of reading status
[0380] The device captures the page of the book the user is looking at with the front camera and sends the information to the server. At the same time, the device continues to acquire gaze information using the eye tracker, and this information is continuously sent to the server.
[0381] 3. Analysis of reading content
[0382] The server receives the text and gaze data and analyzes it. This analysis identifies which part the user is reading and understands what they are reading. Based on this information, the server can select appropriate background music or detect difficult passages.
[0383] 4. Select and play appropriate background music
[0384] The server selects appropriate background music based on the reading content analyzed. The selected background music data is sent to the device, which then uses a bone conduction speaker to provide the background music to the user. This allows music to be played according to the situation, enhancing the reading atmosphere.
[0385] 5. Assistance with difficult-to-read passages
[0386] If the server determines that a particular word or context is difficult for the user to understand, it generates a meaning or summary of that word or context, sends this information to the device, and provides it to the user audibly through a bone conduction speaker. If necessary, it may also be displayed as visual feedback on the feedback panel.
[0387] Specific operation example
[0388] Reading situations
[0389] A user is reading a historical novel and is about to enter a battle scene. The device captures the scene through the front camera and acquires gaze information.
[0390] The server receives the captured text and gaze information and analyzes it to determine that it is a battle scene.
[0391] The server selects epic battle music that heightens the tension and sends it to the device.
[0392] The device provides the music to the user through a bone conduction speaker.
[0393] If the difficult word "Yamato Taro Yoshiie" appears in the middle of a battle scene,
[0394] The server determines that this word is difficult for the user to understand and generates a meaning and background such as "Hachiman Taro Yoshiie was a military commander in the Heian period, and was later enshrined as a god at Hachiman Shrine."
[0395] The device provides this information to the user through a bone conduction speaker.
[0396] This allows users to immerse themselves even more deeply in the world of the book through music and explanations.
[0397] As can be seen, the present invention enhances the reading experience and allows readers to become immersed in the world of the book.
[0398] The processing flow will be explained below.
[0399] Step 1:
[0400] The user puts on the "AI reading glasses." The device automatically starts up and activates the eye tracker to follow the user's gaze and the front camera to capture the text being read.
[0401] Step 2:
[0402] The device uses the front camera to capture the page of the book the user is looking at, formats the captured image data, and saves it in temporary storage.
[0403] Step 3:
[0404] The device uses an eye tracker to acquire user gaze information and temporarily stores the gaze data.
[0405] Step 4:
[0406] The device sends the captured image data and gaze data as a set to the server in real time.
[0407] Step 5:
[0408] The server converts the received image data into text data through an OCR (optical character recognition) process. Based on the converted text data and gaze data, the server performs natural language analysis to identify what the user is reading and where they are reading from.
[0409] Step 6:
[0410] Based on the analyzed text data, the server selects the appropriate background music for the scene. For example, if it's a battle scene, it will select epic battle music.
[0411] Step 7:
[0412] The server sends the selected background music data to the device, which then plays it through the bone conduction speaker, allowing users to enjoy music suitable for reading.
[0413] Step 8:
[0414] The server uses gaze data and text data to detect words and contexts that are difficult to understand, such as difficult historical or academic terms.
[0415] Step 9:
[0416] The server generates meaning and context for each difficult word and context that is detected, which is then converted into audio data.
[0417] Step 10:
[0418] The server sends the generated audio data to the device, which then plays it back through a bone conduction speaker and optionally displays visual information on a feedback panel.
[0419] Through the above steps, the present invention improves the reading experience and provides an environment where users can immerse themselves in the world of books.
[0420] Example 1
[0421] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0422] The modern reading experience remains static, making it difficult for readers to deeply understand the meaning of text or customize their reading environment. Furthermore, there is a lack of ways to quickly understand the meaning of difficult words or contexts when encountering them. Furthermore, there is no function to select appropriate background music, making it difficult to create an environment conducive to immersion in reading.
[0423] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0424] In this invention, the server includes means for detecting the user's gaze direction using a camera installed in front, gaze tracking means for collecting the user's gaze information in real time, machine learning means for analyzing the content being read based on the collected gaze information and data from the camera, background sound selection means for selecting appropriate background sound based on the analyzed content, bone conduction sound means for playing the selected background sound, and explanation generation means for generating and providing explanations of difficult words and contexts as needed, thereby improving the reading experience, enabling the reader to deeply understand the content of the book, and providing an appropriate acoustic environment.
[0425] The "photography device" is a device installed in front of the user to detect the direction of the user's line of sight.
[0426] "Eye tracking means" is a technology for collecting user eye gaze information in real time.
[0427] "Machine learning means" is a technology that analyzes the content of reading based on collected gaze information and data from a camera.
[0428] The "background sound selection means" is a technique for selecting appropriate background sound based on the analyzed content.
[0429] "Bone conduction sound means" is a technology for reproducing selected background sounds.
[0430] "Explanation generator" is a technology for generating and providing explanations of difficult words and contexts as needed.
[0431] The present invention relates to a glasses-type reading assistance device for improving reading experience, which includes a front-mounted camera, an eye tracking unit, a machine learning unit, a background sound selection unit, a bone conduction sound unit, and an explanation generation unit.
[0432] When a user puts on the eyeglass-type reading assistance device, the device automatically starts up and activates the eye tracker (eye tracking means) to track the user's gaze and the front camera (photography device) to capture the text being read, making the device ready for reading.
[0433] Next, the device captures the page of the book the user is looking at with the front camera and sends the information to the server. At the same time, the device continues to acquire gaze information using the eye tracker, which is then continuously sent to the server.
[0434] The server analyzes the transmitted gaze data and captured text images. It uses machine learning methods to identify which part the user is reading and understand what is being read. Based on this information, the server can select appropriate background audio or detect difficult words or contexts.
[0435] The server then selects appropriate background sounds based on the analysis results. The selected background sound data is sent to the device, which then provides the sound to the user using bone conduction audio, thereby playing music appropriate to the situation and enhancing the reading atmosphere.
[0436] Furthermore, if the server determines that a particular word or sentence is difficult for the user to understand as a result of the analysis, it generates a meaning or summary of the word or sentence. This information is sent to the terminal, which then provides it to the user audibly through bone conduction acoustic means. If necessary, the terminal may also display the information on a feedback panel as visual feedback.
[0437] Specific examples
[0438] For example, suppose a user is reading a historical novel. When the user comes to a battle scene in the book, the device captures the scene through the front camera and acquires gaze information. The server receives the captured text and gaze information and analyzes that it is a battle scene. As a result, the server selects epic battle sounds to heighten the tension and transmits them to the device. The device then provides this sound to the user through bone conduction audio.
[0439] Also, if the user comes across a difficult word like "Yoshiie Hachiman" while reading, the server will determine that the word is difficult for the user to understand and generate a meaning and background such as "Yoshiie Hachiman was a military commander in the Heian period, and was later enshrined as a god at Hachiman Shrine." The terminal will provide this information to the user through bone conduction acoustic means.
[0440] Prompt Sentence Examples
[0441] "Please select background music that matches the battle scenes in a historical novel and generate commentary for Yawata Taro Yoshiie."
[0442] As described above, the present invention improves the reading experience, allows the reader to understand the contents of the book more deeply, and provides an appropriate acoustic environment.
[0443] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0444] Step 1:
[0445] The user puts on the "glasses-type reading assistance device." The device automatically starts up, and the eye tracker (a means of tracking gaze) and the front camera (a photographing device) for capturing the text being read are enabled. This allows the device to obtain gaze information in real time, preparing the user's reading environment.
[0446] Step 2:
[0447] The device captures the page of the book the user is looking at through the front camera. At the same time, the eye tracker collects the user's gaze information. As input, the device receives image data and gaze data from the front camera, and sends these data to the server in real time. As output, the captured image data and gaze data are sent to the server.
[0448] Step 3:
[0449] The server analyzes the received gaze data and captured image. It uses gaze data and captured image as input. It uses machine learning algorithms to identify which part the user is reading and performs text analysis. The output is the specific information about the reading part and the analysis results.
[0450] Step 4:
[0451] The server selects appropriate background audio based on the analysis results. It uses the specific information about the reading section and the analysis results as input. The server uses a background audio selection algorithm to select music that matches the analyzed content. The selected background audio data is generated as output.
[0452] Step 5:
[0453] The server sends the selected background sound data to the terminal. The background sound data is used as input. The terminal receives this sound data and provides music to the user using bone conduction sound means. As output, appropriate background music is played for the user.
[0454] Step 6:
[0455] The server identifies difficult words and sentences and generates their meaning and context. It uses gaze data and text analysis results as input. It uses an explanation generation algorithm to create explanations for difficult words and contexts. The output is the generated explanation data.
[0456] Step 7:
[0457] The server sends the generated explanation data to the terminal, which uses the explanation data as input. The terminal receives the explanation data and provides audio explanations to the user through bone conduction acoustic means, and also provides visual feedback if necessary. As output, the user is provided with audio and visual explanations of difficult words and contexts.
[0458] Through the above processing steps, the present invention can improve the reading experience, allow users to understand the contents of the book more deeply, and provide a suitable acoustic environment.
[0459] (Application example 1)
[0460] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0461] Traditional reading experiences have the drawback of making it difficult for readers to gain a deep understanding of the content or experience a sense of realism. In particular, when encountering difficult words or context, there are limited ways to receive on-the-spot explanations, which frequently interrupts the flow of reading. There is also a demand for incorporating background music and sound effects into reading to improve the quality of the reading experience.
[0462] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0463] In this invention, the server includes: means for detecting the reader's gaze direction using a camera installed in front; gaze measurement means for collecting the reader's gaze information in real time; artificial intelligence means for analyzing the content being read based on the collected gaze information and data from the front camera; background music selection means for selecting appropriate background music based on the analyzed content; a sound wave conduction audio device for playing the selected background music; explanation generation means for generating and providing explanations of difficult words and context as needed; data communication means for transmitting the text and gaze information being read to the data server, the server analyzing the transmitted data, and providing the user with the analysis results; and audio control means for transmitting the background music selected based on the analysis results to the smart device and for the smart device to play the background music. This significantly improves the quality of the reading experience, allowing the reader to deeply understand the content and experience a sense of realism.
[0464] 1. "Camera" refers to a device that is installed in front of the user to detect the direction of their gaze and capture images in real time.
[0465] 2. "Gaze measurement means" refers to technology or equipment used to collect readers' gaze information in real time and analyze that information.
[0466] 3. "Artificial intelligence means" refers to algorithms and processes used to analyze the content of a reading based on collected gaze information and camera data.
[0467] 4. "Background music selection means" refers to a method or system for selecting appropriate background music based on the analyzed reading content.
[0468] 5. "Sound wave conduction sound device" means a bone conduction speaker or similar sound reproduction device that provides a user with selected background music.
[0469] 6. "Explanation generation means" refers to technology that generates explanations for difficult words or contexts and provides them to users in audio or visual form.
[0470] 7. "Data communication means" refers to the network technology and protocols used to transmit text and eye gaze information during reading to a data server and receive analysis results.
[0471] 8. "Audio Control Means" means the control technology or algorithm for playing background music selected based on the analysis results on a smart device.
[0472] 9. "Data server" refers to a computer system installed on the cloud or elsewhere that analyzes transmitted data and provides the results.
[0473] To implement this invention, the following system configuration and its operating procedures will be specifically described. The system includes a camera, a gaze measurement means, an artificial intelligence means, a background music selection means, a sound-transmitting audio device, a description generation means, a data communication means, and an audio control means.
[0474] Hardware and Software Use
[0475] Camera: Uses the front-facing camera on your smartphone or tablet to capture where you're looking and the text you're reading.
[0476] Eye tracking method: Use software with eye tracking functionality (e.g., the GazeTracking library).
[0477] Artificial intelligence measures: Machine learning models are used to analyze collected gaze information and captured text data.
[0478] Background music selection method: Use a music library and music recommendation algorithms to select background music appropriate for the reading content.
[0479] Sound wave conduction sound device: Plays background music using bone conduction speakers or smartphone speakers.
[0480] Explanation generation: Use text generation models and text-to-speech technologies (e.g., Google Cloud Text-to-Speech API) to generate explanations for difficult words and contexts and convert them into audio.
[0481] Data communication means: An internet communication protocol (e.g., HTTP request) is used to send text data and gaze data to the cloud server and return the analysis results to the smart device.
[0482] Sound control means: Use the music playback library to play background music on the smart device based on the analysis results sent from the cloud server.
[0483] Specific examples
[0484] Reading situations
[0485] Assume a user is reading a historical novel on their smartphone. When they reach a battle scene in the book, the system works as follows:
[0486] 1. Eye Tracking and Text Capture:
[0487] The page the user is reading is captured through the smartphone camera.
[0488] At the same time, the gaze measurement means collects the user's gaze data in real time.
[0489] 2. Data analysis and background music selection:
[0490] The captured text data and gaze data are transmitted to a cloud server.
[0491] The cloud server uses machine learning models to analyze the data and identify that the user is reading a fight scene.
[0492] Based on the analysis results, epic battle music is selected and the music data is sent back to the smartphone.
[0493] 3. Play background music and explain difficult passages:
[0494] The smartphone plays battle music through a sonic wave transmission audio device.
[0495] If the difficult word "Yahata Taro Yoshiie" appears in the middle of a battle scene, the cloud server will determine that the word is difficult to understand and generate its meaning and background.
[0496] The smartphone plays the generated explanation as audio and provides it to the user.
[0497] Prompt Sentence Examples
[0498] A user is reading a historical novel and is currently approaching a battle scene. Choose epic battle music to match this scene and explain the meaning and background of the difficult passage "Yaman Taro Yoshiie."
[0499] This will significantly improve the quality of the reading experience, allowing readers to gain a deeper understanding of the content and experience a more immersive experience.
[0500] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0501] Step 1:
[0502] The user opens the smartphone, launches the app, and selects an e-book. The input is the e-book file selected by the user and the initial gaze information data. The output is a state where camera capture and gaze tracking are ready. The camera is placed in front and the eye tracking function is enabled.
[0503] Step 2:
[0504] The device uses a front camera to capture text data of the page the user is reading and collects gaze data in real time through an eye gaze measurement means. The input is the camera image and real-time gaze position data. The output is the captured text data and gaze data.
[0505] Step 3:
[0506] The captured text data and gaze data are sent from the terminal to the server. The input is the text data and gaze data sent from the terminal. The output is a notification that the text data and gaze data have been received. These data are sent to the cloud server using a data communication means.
[0507] Step 4:
[0508] The server analyzes the received text data and gaze data using artificial intelligence. The input is the text data and gaze data. The output is the analysis result of what the user is reading. For example, it may be determined that the user is reading a battle scene.
[0509] Step 5:
[0510] The server selects appropriate background music based on the analysis results. The input is the analyzed reading content (e.g., a battle scene). The output is a URL or file of the selected background music data. Using the background music selection means, epic music suitable for the battle scene is selected.
[0511] Step 6:
[0512] The server sends the selected background music data to the terminal. The input is the URL or file of the background music data. The output is the background music data sent to the terminal. The music data is sent to the smartphone using a data communication means.
[0513] Step 7:
[0514] The background music data received by the terminal is played on an acoustic wave conduction acoustic device. The input is the received background music data. The output is the background music being listened to by the user. Using the acoustic control means, the background music is played on the smartphone's bone conduction speaker.
[0515] Step 8:
[0516] The server identifies difficult-to-read passages and generates their meaning and background using an explanation generation means. The input is the analyzed text data and information on the identified difficult-to-read passages. The output is the generated explanatory text. For example, the text generated for "Hachiman Taro Yoshiie" is "He was a military commander in the Heian period, and later enshrined as a god at Hachiman Shrine."
[0517] Step 9:
[0518] The server transmits the explanation generated by the explanation generation means to the terminal. The input is the generated explanation text. The output is the explanation audio data transmitted to the terminal. The explanation data is transmitted to the smartphone using the data communication means.
[0519] Step 10:
[0520] The terminal plays the received explanatory data on an acoustic wave conduction audio device. The input is the received explanatory audio data. The output is the explanatory audio that the user is listening to. Using the audio control means, the explanation is played through a bone conduction speaker.
[0521] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0522] This invention is a glasses-type reading assistance device that combines a front camera, gaze tracking means, artificial intelligence means, background music selection means, bone conduction speakers, explanation generation means, and an emotion engine to improve the reading experience. This device analyzes the content being read based on the user's gaze and emotions, plays background music appropriate to the situation, and provides support for difficult words and context. The program processing of this system is explained below in natural language.
[0523] 1. Setup Phase
[0524] The user puts on the AI reading glasses. The device automatically activates the eye tracker to track the user's gaze and the front camera to capture the text being read. It also activates an emotion engine that analyzes the user's facial expressions and tone of voice.
[0525] 2. Real-time understanding of reading status
[0526] The device captures the page of the book the user is looking at with the front camera and sends the information to the server. At the same time, the device continues to acquire gaze information using the eye tracker. The emotion engine also analyzes the user's facial expressions and tone of voice to acquire emotional data. This information is continuously sent to the server.
[0527] 3. Analysis of reading content
[0528] The server receives the text, gaze data, and emotion data and analyzes them. The server recognizes the text data using OCR (optical character recognition) technology and combines it with the gaze data to identify the specific part the user is reading. It also analyzes the emotion data to determine the user's current emotional state.
[0529] 4. Select and play appropriate background music
[0530] The server selects appropriate background music based on the reading content and the user's emotional data analyzed. For example, if the user is feeling tense while reading, it selects relaxing music, and if the user is enjoying themselves, it selects music that enhances the mood. The selected background music data is sent to the device, which then uses a bone conduction speaker to provide the background music to the user.
[0531] 5. Assistance with difficult-to-read passages
[0532] If the server detects words or contexts that are difficult to read from the gaze data and text data, it generates a meaning and summary of the words or contexts. The generated information is converted into audio data and sent from the server to the device. The device then provides the audio data to the user through a bone conduction speaker. It may also display visual information on a feedback panel as needed.
[0533] Specific operation example
[0534] Reading situations
[0535] A user is reading a mystery novel. The climax of the story is approaching, and the user's heart rate is rising. The device captures the scene through the front camera and acquires gaze and emotion information.
[0536] The server receives the captured text, gaze information, and emotion data and analyzes it to determine whether it is a climax scene.
[0537] The server selects relaxing music to relieve tension based on the user's emotional state and sends it to the terminal.
[0538] The device provides the music to the user through a bone conduction speaker.
[0539] When the difficult name "Albaiohiko" appears in the climax scene,
[0540] The server determines that this name is difficult for the user to understand and generates a meaning and background, such as "Albaiohiko is a mysterious detective and a character with a hidden past."
[0541] The device provides this information to the user through a bone conduction speaker.
[0542] This allows users to receive appropriate music and explanations while reading, allowing them to become more immersed in the world of the book.
[0543] As described above, the present invention improves the reading experience and provides an environment in which users can immerse themselves in the world of books.
[0544] The processing flow will be explained below.
[0545] Step 1:
[0546] The user puts on the "AI reading glasses" and turns on the device, which activates the eye tracker, front-facing camera, and emotion engine.
[0547] Step 2:
[0548] The device uses the front camera to periodically capture the page of the book the user is looking at, and the captured image data is temporarily stored.
[0549] Step 3:
[0550] The device uses an eye tracker to acquire the user's gaze information in real time and temporarily stores this gaze data.
[0551] Step 4:
[0552] The device uses an emotion engine to analyze the user's facial expressions and tone of voice to determine their current emotional state, and emotion data is also stored.
[0553] Step 5:
[0554] The terminal transmits the captured image data, gaze data, and emotion data as a set to the server in real time.
[0555] Step 6:
[0556] The server converts the received image data into text data using OCR (optical character recognition) technology, and then analyzes the text data by combining it with gaze data and emotion data.
[0557] Step 7:
[0558] Based on the analysis, the server identifies the specific part the user is reading and makes a comprehensive assessment of the user's emotional state.
[0559] Step 8:
[0560] The server selects appropriate background music based on the analysis results and emotional data. For example, it selects relaxing music for tense scenes and music that enhances the atmosphere for happy scenes.
[0561] Step 9:
[0562] The server sends the selected background music data to the device, which then uses a bone conduction speaker to provide the music to the user.
[0563] Step 10:
[0564] The server detects difficult words and contexts from gaze data and text data, and generates meaning and background information for the detected content.
[0565] Step 11:
[0566] The server converts the generated explanation data into audio data and sends it to the terminal, which then provides the audio data to the user through a bone conduction speaker. The terminal also displays visual information on a feedback panel as needed.
[0567] As a specific example of how it works, when a user is reading a mystery novel and reaches a tense climax, the emotion engine detects the user's tension.
[0568] The server selects relaxing background music and sends it to the device.
[0569] The device plays the background music through a bone conduction speaker, helping the user relax.
[0570] When the difficult name "Albaiohiko" appears, the server generates background information about the name and sends it to the terminal as voice data.
[0571] The device plays audio to help the user understand.
[0572] In this way, the present invention takes into account the user's emotions while providing appropriate music and information support to enhance the reading experience.
[0573] Example 2
[0574] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0575] While various technologies have been proposed to improve the reading experience, systems that provide real-time support based on changes in the reader's gaze and emotions are not yet fully developed. In particular, technologies that provide appropriate background music while reading and clearly explain difficult words and contexts are in need of further development. This creates a demand for systems that can enhance immersion and comprehension while reading.
[0576] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0577] In this invention, the server includes means for detecting the reader's gaze direction using an image sensor installed in front, a gaze measurement device for collecting the reader's gaze information in real time, artificial intelligence means for analyzing the content being read based on the collected gaze information and data from the front image sensor, background music selection means for selecting appropriate background music based on the analyzed content and the reader's emotional data, a bone conduction audio device for playing the selected background music, explanation generation means for generating and providing explanations of difficult terms and context as necessary, and voice conversion means for converting the analyzed data into voice data and providing it to the reader. This makes it possible to provide dynamic support in real time according to the reader's gaze and emotions, significantly improving the quality of the reading experience.
[0578] A "forward-mounted image sensor" is an image capture device mounted in front of the reader to detect the reader's line of sight.
[0579] The "means for detecting the reader's line of sight" is a mechanism that uses an image sensor installed in front to identify where the reader is looking.
[0580] An "eye gaze measurement device that collects reader gaze information in real time" is a device that instantly tracks the reader's eye movements and collects that data.
[0581] "Artificial intelligence means" refers to technology that analyzes collected gaze information and data from the forward image sensor to understand the content and behavior of the reader while reading.
[0582] The "background music selection means" is a mechanism for selecting optimal background music based on the analyzed content and emotional data.
[0583] A "bone conduction audio device" is an audio playback device that allows selected background music to be heard directly through the bones.
[0584] "Explanation generation means" is a technology that generates the meaning and summary of text to provide easy-to-understand explanations of difficult terms and context.
[0585] "Speech conversion means" refers to a technology that converts the generated commentary and information into audio data and provides it to the reader.
[0586] The present invention is a glasses-type device system designed to enhance the reading experience. The device combines a front-mounted image sensor, an eye tracking device, an artificial intelligence means, a background music selection means, a bone conduction sound device, a commentary generation means, and a voice conversion means to provide optimal reading support to users. Detailed embodiments of the system are described below.
[0587] When a user puts on the AI reading glasses, the device automatically starts up. The device includes an eye tracker to track the user's gaze and a forward-facing image sensor to capture the text being read. An emotion engine then starts up, analyzing the user's facial expressions and tone of voice.
[0588] The device captures the page of the book the user is looking at with a front image sensor and sends the information to the server. At the same time, it acquires gaze information using an eye tracker and collects emotion data with an emotion engine. This information is continuously sent to the server.
[0589] The server converts the image data into text using OCR (optical character recognition) technology, such as software like Tesseract OCR. The acquired text data is combined with gaze data to identify the user's current reading position. At the same time, a sentiment analysis algorithm is used to analyze the emotion data and determine the user's current emotional state.
[0590] The server selects appropriate background music based on the analyzed content. For example, if you are feeling tense while reading, it will select relaxing music, and if you are enjoying yourself, it will select music that will enhance the atmosphere. The background music data is sent to the terminal and provided to the user via a bone conduction audio device.
[0591] Additionally, if the server detects words or contexts that are difficult to understand from the gaze data and text data, it generates a meaning or summary of those words or contexts. This content is generated using a generative AI model (such as the GPT series). The generated information is converted into audio data and sent from the server to the device. The device then provides the audio data to the user through a bone conduction audio device. If necessary, the information is also displayed on a visual feedback panel.
[0592] As a concrete example, consider a case where a user is reading a mystery fiction book. As the climax approaches, the user's heart rate increases. The device captures this scene through the forward-facing image sensor and acquires gaze information and emotional information.
[0593] The server receives the captured text, gaze information, and emotion data and analyzes it to determine whether it is a climax scene.
[0594] The server selects relaxing music to relieve tension based on the user's emotional state and sends it to the terminal.
[0595] The terminal provides the music to the user through a bone conduction audio device.
[0596] If a mysterious character called "Detective Albion" appears in the climax scene,
[0597] The server determines that this character is difficult for the user to understand, and uses a generative AI model to generate meaning and background, such as "Detective Albaio is a mysterious detective and a character with a hidden past."
[0598] The terminal provides this information to the user through a bone conduction audio device.
[0599] This system allows users to receive appropriate music and explanations while reading, allowing them to become more immersed in the world of the book. In addition, by linking it with a generative AI model, it is possible to generate prompts to aid the reader in comprehension.
[0600] Prompt Sentence Examples
[0601] Please provide a brief background on the character "Detective Albio."
[0602] Choose relaxing music to listen to while you read.
[0603] The above is an embodiment of the present invention, which allows users to have a more comfortable and immersive reading experience.
[0604] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0605] Step 1:
[0606] The user puts on the AI reading glasses. The device automatically starts up, activating the eye tracker to track the user's gaze and the forward-facing image sensor to capture the text being read. The emotion engine also starts up.
[0607] Specifically, the device starts up the image sensor and eye tracker, and goes through an initialization process to prepare for acquiring gaze and video data.
[0608] Input: The user puts on the device.
[0609] Output: Eye tracking and image acquisition ready.
[0610] Step 2:
[0611] The device uses a forward-facing image sensor to capture the page of the book the user is reading, an eye tracker to collect gaze information in real time, and an emotion engine to extract emotional data from facial expressions and tone of voice.
[0612] Specifically, the device continuously captures images of the page using a front image sensor, acquires user gaze data using an eye tracker, and performs facial expression and voice analysis using an emotion engine.
[0613] Input: User's gaze information, forward image, emotion data.
[0614] Output: Captured image data, gaze data, emotion data.
[0615] Step 3:
[0616] The device continuously transmits captured image data, gaze data, and emotion data to the server.
[0617] Specifically, the terminal compresses the collected data at regular intervals and transmits it to the server via packet communication. Since the data is transmitted in real time, network bandwidth and communication efficiency are taken into consideration.
[0618] Input: Captured image data, gaze data, emotion data.
[0619] Output: Various data sent to the server.
[0620] Step 4:
[0621] The server converts the received image data into text using OCR technology (e.g., Tesseract OCR), combines it with gaze data to identify the part the user is reading, and analyzes emotion data to determine the user's emotional state.
[0622] Specifically, the server converts image data into text using an OCR engine, maps it with gaze data, and then uses an emotion analysis algorithm to identify the user's emotional state.
[0623] Input: Transmitted image data, gaze data, emotion data.
[0624] Output: Transformed text data, identified reading passages, and the user's emotional state.
[0625] Step 5:
[0626] The server selects appropriate background music based on the analysis results. For example, if the user is feeling tense, it selects relaxing music, and if the user is having fun, it selects music that will liven up the atmosphere. The selected background music data is sent to the device.
[0627] Specifically, the server selects the most suitable music from multiple music databases, compresses the music data, and sends it to the terminal.
[0628] Input: Parsed reading content, user's emotional state.
[0629] Output: Selected background music data, sent to the device.
[0630] Step 6:
[0631] The background music data received by the terminal is played through a bone conduction audio device.
[0632] Specifically, the device stores background music data in a buffer and plays it in real time through a bone conduction audio device.
[0633] Input: BGM data sent from the server.
[0634] Output: Providing background music to the user.
[0635] Step 7:
[0636] The server uses gaze data and text data to identify words and contexts that the user has difficulty comprehending, and generates meanings and summaries of those words and contexts. The generated information is converted into audio data and sent to the device.
[0637] Specifically, the server analyzes gaze data and text data, generates meaning and summaries using a generative AI model (e.g., GPT-3), and converts this text data into audio data.
[0638] Input: gaze data, text data.
[0639] Output: Generated description data and audio data.
[0640] Step 8:
[0641] The audio data received by the terminal is provided to the user through a bone conduction audio device.
[0642] Specifically, the device decodes the audio data, plays the audio through a bone conduction audio device, and optionally displays visual information on the feedback panel.
[0643] Input: Audio data sent from the server.
[0644] Output: Providing audio information to the user.
[0645] In this way, through the specific processing performed at each step, the user can receive appropriate reading assistance in real time.
[0646] (Application example 2)
[0647] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0648] Conventional reading assistance devices and electronic payment systems are unable to reflect the user's emotional state or gaze data in real time, limiting their ability to provide diverse information and improve the user experience. Furthermore, when shopping online, users often feel stressed when choosing the most suitable product and payment method, which can reduce their motivation to purchase. There is a need to solve these issues and provide users with a richer experience and more efficient shopping support.
[0649] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0650] In this invention, the server includes: means for detecting the user's gaze direction using a camera installed in front; gaze measurement means for collecting user gaze information in real time; artificial intelligence means for analyzing the content of reading and online activities based on the collected gaze information and data from the front camera; background music selection means for selecting appropriate background music based on the analyzed content; bone conduction speaker means for playing the selected background music; explanation generation means for generating and providing explanations of difficult words and contexts as necessary; product recommendation means for capturing the product page the user is viewing while shopping online and recommending appropriate products and services based on that information and emotional data; and payment assistance means for suggesting the optimal payment method based on the analyzed emotional data and supporting purchasing behavior. This enables optimal music selection and product recommendation based on emotional analysis.
[0651] A "front camera" is a photographic device installed on the front of a device to capture the user's line of sight and what they are viewing.
[0652] "Gaze measurement means" refers to technology or equipment that collects a user's gaze information in real time and analyzes its direction and focus.
[0653] "Artificial intelligence means" refers to technologies and systems that analyze collected gaze information and data from the front camera to understand the content of reading and online activities.
[0654] The "background music selection means" refers to a technology or system for selecting optimal background music based on the analyzed content.
[0655] "Bone conduction speaker means" means an audio device that uses bone conduction technology to transmit selected background music to a user.
[0656] "Explanation generation means" refers to technology or systems that generate explanations of difficult words or contexts as needed and provide them to users.
[0657] A "product recommendation method" is a technology or system that captures the product page a user is viewing while shopping online and recommends appropriate products or services based on that information and emotional data.
[0658] "Payment assistance methods" are technologies and systems that suggest the most suitable payment method to users based on analyzed emotional data and assist them in their purchasing behavior.
[0659] The present invention is implemented as a system for improving the reading experience and online shopping experience by a device that combines a forward-facing camera, a gaze measurement means, an artificial intelligence means, a background music selection means, a bone conduction speaker means, a description generation means, a product recommendation means, and a payment assistance means.
[0660] 1. Hardware configuration:
[0661] The front-facing camera is located at the front of the device to capture the user's line of sight and what they are viewing.
[0662] The gaze measurement means is used to collect user gaze information in real time.
[0663] The bone conduction speaker is used to convey selected background music and explanatory information to the user.
[0664] 2. Software configuration:
[0665] The artificial intelligence means analyzes the collected gaze information and data from the front camera to understand the content of your reading and online activities.
[0666] The background music selection means selects appropriate background music based on the analyzed content.
[0667] The explanation generating means generates explanations of difficult words and contexts as needed and provides them to the user.
[0668] The product recommendation means captures the product page that the user is viewing while shopping online and recommends appropriate products and services based on that information and emotional data.
[0669] The payment assistance tool suggests the most suitable payment method to the user based on the analyzed emotional data, and supports purchasing behavior.
[0670] 3. Data processing and calculation:
[0671] The video data acquired from the front camera is sent to a server and converted into text data using OCR technology.
[0672] The gaze data and emotion analysis data acquired by the gaze measurement means are analyzed by the artificial intelligence means to identify the specific part the user is reading.
[0673] The background music selection means selects the most suitable music based on the emotion and theme of the analyzed text and plays it through a bone conduction speaker.
[0674] The product recommendation and payment assistance methods capture the product page the user is viewing while shopping online, and suggest the most suitable products and payment methods based on that information and emotional data.
[0675] 4. Specific examples:
[0676] For example, if a user is reading a mystery novel, emotion analysis can tell if the user is feeling tense. The system can detect this and play relaxing music through a bone conduction speaker. It can also generate and provide audio explanations of difficult words and context that appear in the climax scene.
[0677] 5. Example of a generative AI model and prompt:
[0678] Example prompt sentence:
[0679] Capture the product page the user is viewing with a camera, collect gaze and emotion data, and send that data to your server to recommend the best products and payment methods.
[0680] This allows users to receive optimal music and descriptions, product recommendations, and payment assistance while reading or shopping online, greatly improving their experience.
[0681] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0682] Step 1:
[0683] The user puts on the device ("Smart Payment Assist Glasses") and the system starts automatically. The input is the user's wearing information. The device activates the camera and gaze measurement means and begins to capture the user's gaze information and facial expressions in real time. The output obtained is gaze data and facial expression data captured in real time.
[0684] Step 2:
[0685] The device uses the front camera to capture an image of the product page the user is viewing. The input is the product page the user is viewing. The image data captured by the camera is sent to the server. The output is the image data sent to the server.
[0686] Step 3:
[0687] The server analyzes the received image data using OCR (Optical Character Recognition) technology to extract text data related to the product. The input is image data captured by a camera. Through OCR analysis, text data is extracted to obtain the specific product information the user is viewing. The output is the extracted text data.
[0688] Step 4:
[0689] The device uses gaze measurement means to collect the user's gaze information in real time. The input is the user's eye movement and focus. The gaze data is continuously sent to a server and used to identify the specific location the user is looking at. The output is the identified gaze information.
[0690] Step 5:
[0691] The server uses an emotion engine to analyze the user's facial expressions and tone of voice to obtain current emotion data. The inputs are the user's facial expressions and tone of voice. The emotion data is analyzed along with the gaze data to help identify the user's current emotional state. The output is the identified emotion data.
[0692] Step 6:
[0693] The server comprehensively analyzes gaze data, text data, and emotion data to determine the products in which the user is interested and the degree of their willingness to purchase them. The inputs are gaze data, text data, and emotion data. The analysis results in the recommendation of appropriate products and services. The output is information about recommended products and services.
[0694] Step 7:
[0695] Based on the analyzed content, the server selects appropriate background music using a background music selection means. The inputs are the user's emotional state and the text content being viewed. The selected music is sent to the terminal and played to the user through a bone conduction speaker. The resulting output is the selected background music.
[0696] Step 8:
[0697] The server generates explanations of difficult words and contexts for the user and converts them into audio data. The input is the text content the user is viewing and the identified difficult parts. The explanation generation means generates a summary and sends it to the terminal as audio data. The output obtained is explanatory information as audio data.
[0698] Step 9:
[0699] The server proposes the optimal payment method based on the analyzed emotional data and purchasing intention information and notifies the user. The input is the user's emotional state and recommended product information. The payment assistance means selects the appropriate payment method and provides assistance to the user. The output obtained is the proposed payment method.
[0700] This allows users to receive the best music and explanations, product recommendations, and payment assistance while reading or shopping online.
[0701] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0702] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0703] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0704] [Third embodiment]
[0705] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0706] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0707] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0708] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0709] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0710] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0711] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0712] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0713] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0714] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0715] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0716] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0717] The present invention provides a glasses-type reading assistance device for improving the reading experience. The device includes a front camera, a gaze measurement unit, an artificial intelligence unit, a background music selection unit, a bone conduction speaker, and a description generation unit. The program processing of this system is explained below in natural language.
[0718] 1. Setup Phase
[0719] The user puts on the "AI reading glasses." The device automatically starts up and activates the eye tracker to follow the user's gaze and the front camera to capture the text being read. The device is then ready for reading.
[0720] 2. Real-time understanding of reading status
[0721] The device captures the page of the book the user is looking at with the front camera and sends the information to the server. At the same time, the device continues to acquire gaze information using the eye tracker, and this information is continuously sent to the server.
[0722] 3. Analysis of reading content
[0723] The server receives the text and gaze data and analyzes it. This analysis identifies which part the user is reading and understands what they are reading. Based on this information, the server can select appropriate background music or detect difficult passages.
[0724] 4. Select and play appropriate background music
[0725] The server selects appropriate background music based on the reading content analyzed. The selected background music data is sent to the device, which then uses a bone conduction speaker to provide the background music to the user. This allows music to be played according to the situation, enhancing the reading atmosphere.
[0726] 5. Assistance with difficult-to-read passages
[0727] If the server determines that a particular word or context is difficult for the user to understand, it generates a meaning or summary of that word or context, sends this information to the device, and provides it to the user audibly through a bone conduction speaker. If necessary, it may also be displayed as visual feedback on the feedback panel.
[0728] Specific operation example
[0729] Reading situations
[0730] A user is reading a historical novel and is about to enter a battle scene. The device captures the scene through the front camera and acquires gaze information.
[0731] The server receives the captured text and gaze information and analyzes it to determine that it is a battle scene.
[0732] The server selects epic battle music that heightens the tension and sends it to the device.
[0733] The device provides the music to the user through a bone conduction speaker.
[0734] If the difficult word "Yamato Taro Yoshiie" appears in the middle of a battle scene,
[0735] The server determines that this word is difficult for the user to understand and generates a meaning and background such as "Hachiman Taro Yoshiie was a military commander in the Heian period, and was later enshrined as a god at Hachiman Shrine."
[0736] The device provides this information to the user through a bone conduction speaker.
[0737] This allows users to immerse themselves even more deeply in the world of the book through music and explanations.
[0738] As can be seen, the present invention enhances the reading experience and allows readers to become immersed in the world of the book.
[0739] The processing flow will be explained below.
[0740] Step 1:
[0741] The user puts on the "AI reading glasses." The device automatically starts up and activates the eye tracker to follow the user's gaze and the front camera to capture the text being read.
[0742] Step 2:
[0743] The device uses the front camera to capture the page of the book the user is looking at, formats the captured image data, and saves it in temporary storage.
[0744] Step 3:
[0745] The device uses an eye tracker to acquire user gaze information and temporarily stores the gaze data.
[0746] Step 4:
[0747] The device sends the captured image data and gaze data as a set to the server in real time.
[0748] Step 5:
[0749] The server converts the received image data into text data through an OCR (optical character recognition) process. Based on the converted text data and gaze data, the server performs natural language analysis to identify what the user is reading and where they are reading from.
[0750] Step 6:
[0751] Based on the analyzed text data, the server selects the appropriate background music for the scene. For example, if it's a battle scene, it will select epic battle music.
[0752] Step 7:
[0753] The server sends the selected background music data to the device, which then plays it through the bone conduction speaker, allowing users to enjoy music suitable for reading.
[0754] Step 8:
[0755] The server uses gaze data and text data to detect words and contexts that are difficult to understand, such as difficult historical or academic terms.
[0756] Step 9:
[0757] The server generates meaning and context for each difficult word and context that is detected, which is then converted into audio data.
[0758] Step 10:
[0759] The server sends the generated audio data to the device, which then plays it back through a bone conduction speaker and optionally displays visual information on a feedback panel.
[0760] Through the above steps, the present invention improves the reading experience and provides an environment where users can immerse themselves in the world of books.
[0761] Example 1
[0762] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0763] The modern reading experience remains static, making it difficult for readers to deeply understand the meaning of text or customize their reading environment. Furthermore, there is a lack of ways to quickly understand the meaning of difficult words or contexts when encountering them. Furthermore, there is no function to select appropriate background music, making it difficult to create an environment conducive to immersion in reading.
[0764] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0765] In this invention, the server includes means for detecting the user's gaze direction using a camera installed in front, gaze tracking means for collecting the user's gaze information in real time, machine learning means for analyzing the content being read based on the collected gaze information and data from the camera, background sound selection means for selecting appropriate background sound based on the analyzed content, bone conduction sound means for playing the selected background sound, and explanation generation means for generating and providing explanations of difficult words and contexts as needed, thereby improving the reading experience, enabling the reader to deeply understand the content of the book, and providing an appropriate acoustic environment.
[0766] The "photography device" is a device installed in front of the user to detect the direction of the user's line of sight.
[0767] "Eye tracking means" is a technology for collecting user eye gaze information in real time.
[0768] "Machine learning means" is a technology that analyzes the content of reading based on collected gaze information and data from a camera.
[0769] The "background sound selection means" is a technique for selecting appropriate background sound based on the analyzed content.
[0770] "Bone conduction sound means" is a technology for reproducing selected background sounds.
[0771] "Explanation generator" is a technology for generating and providing explanations of difficult words and contexts as needed.
[0772] The present invention relates to a glasses-type reading assistance device for improving reading experience, which includes a front-mounted camera, an eye tracking unit, a machine learning unit, a background sound selection unit, a bone conduction sound unit, and an explanation generation unit.
[0773] When a user puts on the eyeglass-type reading assistance device, the device automatically starts up and activates the eye tracker (eye tracking means) to track the user's gaze and the front camera (photography device) to capture the text being read, making the device ready for reading.
[0774] Next, the device captures the page of the book the user is looking at with the front camera and sends the information to the server. At the same time, the device continues to acquire gaze information using the eye tracker, which is then continuously sent to the server.
[0775] The server analyzes the transmitted gaze data and captured text images. It uses machine learning methods to identify which part the user is reading and understand what is being read. Based on this information, the server can select appropriate background audio or detect difficult words or contexts.
[0776] The server then selects appropriate background sounds based on the analysis results. The selected background sound data is sent to the device, which then provides the sound to the user using bone conduction audio, thereby playing music appropriate to the situation and enhancing the reading atmosphere.
[0777] Furthermore, if the server determines that a particular word or sentence is difficult for the user to understand as a result of the analysis, it generates a meaning or summary of the word or sentence. This information is sent to the terminal, which then provides it to the user audibly through bone conduction acoustic means. If necessary, the terminal may also display the information on a feedback panel as visual feedback.
[0778] Specific examples
[0779] For example, suppose a user is reading a historical novel. When the user comes to a battle scene in the book, the device captures the scene through the front camera and acquires gaze information. The server receives the captured text and gaze information and analyzes that it is a battle scene. As a result, the server selects epic battle sounds to heighten the tension and transmits them to the device. The device then provides this sound to the user through bone conduction audio.
[0780] Also, if the user comes across a difficult word like "Yoshiie Hachiman" while reading, the server will determine that the word is difficult for the user to understand and generate a meaning and background such as "Yoshiie Hachiman was a military commander in the Heian period, and was later enshrined as a god at Hachiman Shrine." The terminal will provide this information to the user through bone conduction acoustic means.
[0781] Prompt Sentence Examples
[0782] "Please select background music that matches the battle scenes in a historical novel and generate commentary for Yawata Taro Yoshiie."
[0783] As described above, the present invention improves the reading experience, allows the reader to understand the contents of the book more deeply, and provides an appropriate acoustic environment.
[0784] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0785] Step 1:
[0786] The user puts on the "glasses-type reading assistance device." The device automatically starts up, and the eye tracker (a means of tracking gaze) and the front camera (a photographing device) for capturing the text being read are enabled. This allows the device to obtain gaze information in real time, preparing the user's reading environment.
[0787] Step 2:
[0788] The device captures the page of the book the user is looking at through the front camera. At the same time, the eye tracker collects the user's gaze information. As input, the device receives image data and gaze data from the front camera, and sends these data to the server in real time. As output, the captured image data and gaze data are sent to the server.
[0789] Step 3:
[0790] The server analyzes the received gaze data and captured image. It uses gaze data and captured image as input. It uses machine learning algorithms to identify which part the user is reading and performs text analysis. The output is the specific information about the reading part and the analysis results.
[0791] Step 4:
[0792] The server selects appropriate background audio based on the analysis results. It uses the specific information about the reading section and the analysis results as input. The server uses a background audio selection algorithm to select music that matches the analyzed content. The selected background audio data is generated as output.
[0793] Step 5:
[0794] The server sends the selected background sound data to the terminal. The background sound data is used as input. The terminal receives this sound data and provides music to the user using bone conduction sound means. As output, appropriate background music is played for the user.
[0795] Step 6:
[0796] The server identifies difficult words and sentences and generates their meaning and context. It uses gaze data and text analysis results as input. It uses an explanation generation algorithm to create explanations for difficult words and contexts. The output is the generated explanation data.
[0797] Step 7:
[0798] The server sends the generated explanation data to the terminal, which uses the explanation data as input. The terminal receives the explanation data and provides audio explanations to the user through bone conduction acoustic means, and also provides visual feedback if necessary. As output, the user is provided with audio and visual explanations of difficult words and contexts.
[0799] Through the above processing steps, the present invention can improve the reading experience, allow users to understand the contents of the book more deeply, and provide a suitable acoustic environment.
[0800] (Application example 1)
[0801] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0802] Traditional reading experiences have the drawback of making it difficult for readers to gain a deep understanding of the content or experience a sense of realism. In particular, when encountering difficult words or context, there are limited ways to receive on-the-spot explanations, which frequently interrupts the flow of reading. There is also a demand for incorporating background music and sound effects into reading to improve the quality of the reading experience.
[0803] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0804] In this invention, the server includes: means for detecting the reader's gaze direction using a camera installed in front; gaze measurement means for collecting the reader's gaze information in real time; artificial intelligence means for analyzing the content being read based on the collected gaze information and data from the front camera; background music selection means for selecting appropriate background music based on the analyzed content; a sound wave conduction audio device for playing the selected background music; explanation generation means for generating and providing explanations of difficult words and context as needed; data communication means for transmitting the text and gaze information being read to the data server, the server analyzing the transmitted data, and providing the user with the analysis results; and audio control means for transmitting the background music selected based on the analysis results to the smart device and for the smart device to play the background music. This significantly improves the quality of the reading experience, allowing the reader to deeply understand the content and experience a sense of realism.
[0805] 1. "Camera" refers to a device that is installed in front of the user to detect the direction of their gaze and capture images in real time.
[0806] 2. "Gaze measurement means" refers to technology or equipment used to collect readers' gaze information in real time and analyze that information.
[0807] 3. "Artificial intelligence means" refers to algorithms and processes used to analyze the content of a reading based on collected gaze information and camera data.
[0808] 4. "Background music selection means" refers to a method or system for selecting appropriate background music based on the analyzed reading content.
[0809] 5. "Sound wave conduction sound device" means a bone conduction speaker or similar sound reproduction device that provides a user with selected background music.
[0810] 6. "Explanation generation means" refers to technology that generates explanations for difficult words or contexts and provides them to users in audio or visual form.
[0811] 7. "Data communication means" refers to the network technology and protocols used to transmit text and eye gaze information during reading to a data server and receive analysis results.
[0812] 8. "Audio Control Means" means the control technology or algorithm for playing background music selected based on the analysis results on a smart device.
[0813] 9. "Data server" refers to a computer system installed on the cloud or elsewhere that analyzes transmitted data and provides the results.
[0814] To implement this invention, the following system configuration and its operating procedures will be specifically described. The system includes a camera, a gaze measurement means, an artificial intelligence means, a background music selection means, a sound-transmitting audio device, a description generation means, a data communication means, and an audio control means.
[0815] Hardware and Software Use
[0816] Camera: Uses the front-facing camera on your smartphone or tablet to capture where you're looking and the text you're reading.
[0817] Eye tracking method: Use software with eye tracking functionality (e.g., the GazeTracking library).
[0818] Artificial intelligence measures: Machine learning models are used to analyze collected gaze information and captured text data.
[0819] Background music selection method: Use a music library and music recommendation algorithms to select background music appropriate for the reading content.
[0820] Sound wave conduction sound device: Plays background music using bone conduction speakers or smartphone speakers.
[0821] Explanation generation: Use text generation models and text-to-speech technologies (e.g., Google Cloud Text-to-Speech API) to generate explanations for difficult words and contexts and convert them into audio.
[0822] Data communication means: An internet communication protocol (e.g., HTTP request) is used to send text data and gaze data to the cloud server and return the analysis results to the smart device.
[0823] Sound control means: Use the music playback library to play background music on the smart device based on the analysis results sent from the cloud server.
[0824] Specific examples
[0825] Reading situations
[0826] Assume a user is reading a historical novel on their smartphone. When they reach a battle scene in the book, the system works as follows:
[0827] 1. Eye Tracking and Text Capture:
[0828] The page the user is reading is captured through the smartphone camera.
[0829] At the same time, the gaze measurement means collects the user's gaze data in real time.
[0830] 2. Data analysis and background music selection:
[0831] The captured text data and gaze data are transmitted to a cloud server.
[0832] The cloud server uses machine learning models to analyze the data and identify that the user is reading a fight scene.
[0833] Based on the analysis results, epic battle music is selected and the music data is sent back to the smartphone.
[0834] 3. Play background music and explain difficult passages:
[0835] The smartphone plays battle music through a sonic wave transmission audio device.
[0836] If the difficult word "Yahata Taro Yoshiie" appears in the middle of a battle scene, the cloud server will determine that the word is difficult to understand and generate its meaning and background.
[0837] The smartphone plays the generated explanation as audio and provides it to the user.
[0838] Prompt Sentence Examples
[0839] A user is reading a historical novel and is currently approaching a battle scene. Choose epic battle music to match this scene and explain the meaning and background of the difficult passage "Yaman Taro Yoshiie."
[0840] This will significantly improve the quality of the reading experience, allowing readers to gain a deeper understanding of the content and experience a more immersive experience.
[0841] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0842] Step 1:
[0843] The user opens the smartphone, launches the app, and selects an e-book. The input is the e-book file selected by the user and the initial gaze information data. The output is a state where camera capture and gaze tracking are ready. The camera is placed in front and the eye tracking function is enabled.
[0844] Step 2:
[0845] The device uses a front camera to capture text data of the page the user is reading and collects gaze data in real time through an eye gaze measurement means. The input is the camera image and real-time gaze position data. The output is the captured text data and gaze data.
[0846] Step 3:
[0847] The captured text data and gaze data are sent from the terminal to the server. The input is the text data and gaze data sent from the terminal. The output is a notification that the text data and gaze data have been received. These data are sent to the cloud server using a data communication means.
[0848] Step 4:
[0849] The server analyzes the received text data and gaze data using artificial intelligence. The input is the text data and gaze data. The output is the analysis result of what the user is reading. For example, it may be determined that the user is reading a battle scene.
[0850] Step 5:
[0851] The server selects appropriate background music based on the analysis results. The input is the analyzed reading content (e.g., a battle scene). The output is a URL or file of the selected background music data. Using the background music selection means, epic music suitable for the battle scene is selected.
[0852] Step 6:
[0853] The server sends the selected background music data to the terminal. The input is the URL or file of the background music data. The output is the background music data sent to the terminal. The music data is sent to the smartphone using a data communication means.
[0854] Step 7:
[0855] The background music data received by the terminal is played on an acoustic wave conduction acoustic device. The input is the received background music data. The output is the background music being listened to by the user. Using the acoustic control means, the background music is played on the smartphone's bone conduction speaker.
[0856] Step 8:
[0857] The server identifies difficult-to-read passages and generates their meaning and background using an explanation generation means. The input is the analyzed text data and information on the identified difficult-to-read passages. The output is the generated explanatory text. For example, the text generated for "Hachiman Taro Yoshiie" is "He was a military commander in the Heian period, and later enshrined as a god at Hachiman Shrine."
[0858] Step 9:
[0859] The server transmits the explanation generated by the explanation generation means to the terminal. The input is the generated explanation text. The output is the explanation audio data transmitted to the terminal. The explanation data is transmitted to the smartphone using the data communication means.
[0860] Step 10:
[0861] The terminal plays the received explanatory data on an acoustic wave conduction audio device. The input is the received explanatory audio data. The output is the explanatory audio that the user is listening to. Using the audio control means, the explanation is played through a bone conduction speaker.
[0862] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0863] This invention is a glasses-type reading assistance device that combines a front camera, gaze tracking means, artificial intelligence means, background music selection means, bone conduction speakers, explanation generation means, and an emotion engine to improve the reading experience. This device analyzes the content being read based on the user's gaze and emotions, plays background music appropriate to the situation, and provides support for difficult words and context. The program processing of this system is explained below in natural language.
[0864] 1. Setup Phase
[0865] The user puts on the AI reading glasses. The device automatically activates the eye tracker to track the user's gaze and the front camera to capture the text being read. It also activates an emotion engine that analyzes the user's facial expressions and tone of voice.
[0866] 2. Real-time understanding of reading status
[0867] The device captures the page of the book the user is looking at with the front camera and sends the information to the server. At the same time, the device continues to acquire gaze information using the eye tracker. The emotion engine also analyzes the user's facial expressions and tone of voice to acquire emotional data. This information is continuously sent to the server.
[0868] 3. Analysis of reading content
[0869] The server receives the text, gaze data, and emotion data and analyzes them. The server recognizes the text data using OCR (optical character recognition) technology and combines it with the gaze data to identify the specific part the user is reading. It also analyzes the emotion data to determine the user's current emotional state.
[0870] 4. Select and play appropriate background music
[0871] The server selects appropriate background music based on the reading content and the user's emotional data analyzed. For example, if the user is feeling tense while reading, it selects relaxing music, and if the user is enjoying themselves, it selects music that enhances the mood. The selected background music data is sent to the device, which then uses a bone conduction speaker to provide the background music to the user.
[0872] 5. Assistance with difficult-to-read passages
[0873] If the server detects words or contexts that are difficult to read from the gaze data and text data, it generates a meaning and summary of the words or contexts. The generated information is converted into audio data and sent from the server to the device. The device then provides the audio data to the user through a bone conduction speaker. It may also display visual information on a feedback panel as needed.
[0874] Specific operation example
[0875] Reading situations
[0876] A user is reading a mystery novel. The climax of the story is approaching, and the user's heart rate is rising. The device captures the scene through the front camera and acquires gaze and emotion information.
[0877] The server receives the captured text, gaze information, and emotion data and analyzes it to determine whether it is a climax scene.
[0878] The server selects relaxing music to relieve tension based on the user's emotional state and sends it to the terminal.
[0879] The device provides the music to the user through a bone conduction speaker.
[0880] When the difficult name "Albaiohiko" appears in the climax scene,
[0881] The server determines that this name is difficult for the user to understand and generates a meaning and background, such as "Albaiohiko is a mysterious detective and a character with a hidden past."
[0882] The device provides this information to the user through a bone conduction speaker.
[0883] This allows users to receive appropriate music and explanations while reading, allowing them to become more immersed in the world of the book.
[0884] As described above, the present invention improves the reading experience and provides an environment in which users can immerse themselves in the world of books.
[0885] The processing flow will be explained below.
[0886] Step 1:
[0887] The user puts on the "AI reading glasses" and turns on the device, which activates the eye tracker, front-facing camera, and emotion engine.
[0888] Step 2:
[0889] The device uses the front camera to periodically capture the page of the book the user is looking at, and the captured image data is temporarily stored.
[0890] Step 3:
[0891] The device uses an eye tracker to acquire the user's gaze information in real time and temporarily stores this gaze data.
[0892] Step 4:
[0893] The device uses an emotion engine to analyze the user's facial expressions and tone of voice to determine their current emotional state, and emotion data is also stored.
[0894] Step 5:
[0895] The terminal transmits the captured image data, gaze data, and emotion data as a set to the server in real time.
[0896] Step 6:
[0897] The server converts the received image data into text data using OCR (optical character recognition) technology, and then analyzes the text data by combining it with gaze data and emotion data.
[0898] Step 7:
[0899] Based on the analysis, the server identifies the specific part the user is reading and makes a comprehensive assessment of the user's emotional state.
[0900] Step 8:
[0901] The server selects appropriate background music based on the analysis results and emotional data. For example, it selects relaxing music for tense scenes and music that enhances the atmosphere for happy scenes.
[0902] Step 9:
[0903] The server sends the selected background music data to the device, which then uses a bone conduction speaker to provide the music to the user.
[0904] Step 10:
[0905] The server detects difficult words and contexts from gaze data and text data, and generates meaning and background information for the detected content.
[0906] Step 11:
[0907] The server converts the generated explanation data into audio data and sends it to the terminal, which then provides the audio data to the user through a bone conduction speaker. The terminal also displays visual information on a feedback panel as needed.
[0908] As a specific example of how it works, when a user is reading a mystery novel and reaches a tense climax, the emotion engine detects the user's tension.
[0909] The server selects relaxing background music and sends it to the device.
[0910] The device plays the background music through a bone conduction speaker, helping the user relax.
[0911] When the difficult name "Albaiohiko" appears, the server generates background information about the name and sends it to the terminal as voice data.
[0912] The device plays audio to help the user understand.
[0913] In this way, the present invention takes into account the user's emotions while providing appropriate music and information support to enhance the reading experience.
[0914] Example 2
[0915] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0916] While various technologies have been proposed to improve the reading experience, systems that provide real-time support based on changes in the reader's gaze and emotions are not yet fully developed. In particular, technologies that provide appropriate background music while reading and clearly explain difficult words and contexts are in need of further development. This creates a demand for systems that can enhance immersion and comprehension while reading.
[0917] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0918] In this invention, the server includes means for detecting the reader's gaze direction using an image sensor installed in front, a gaze measurement device for collecting the reader's gaze information in real time, artificial intelligence means for analyzing the content being read based on the collected gaze information and data from the front image sensor, background music selection means for selecting appropriate background music based on the analyzed content and the reader's emotional data, a bone conduction audio device for playing the selected background music, explanation generation means for generating and providing explanations of difficult terms and context as necessary, and voice conversion means for converting the analyzed data into voice data and providing it to the reader. This makes it possible to provide dynamic support in real time according to the reader's gaze and emotions, significantly improving the quality of the reading experience.
[0919] A "forward-mounted image sensor" is an image capture device mounted in front of the reader to detect the reader's line of sight.
[0920] The "means for detecting the reader's line of sight" is a mechanism that uses an image sensor installed in front to identify where the reader is looking.
[0921] An "eye gaze measurement device that collects reader gaze information in real time" is a device that instantly tracks the reader's eye movements and collects that data.
[0922] "Artificial intelligence means" refers to technology that analyzes collected gaze information and data from the forward image sensor to understand the content and behavior of the reader while reading.
[0923] The "background music selection means" is a mechanism for selecting optimal background music based on the analyzed content and emotional data.
[0924] A "bone conduction audio device" is an audio playback device that allows selected background music to be heard directly through the bones.
[0925] "Explanation generation means" is a technology that generates the meaning and summary of text to provide easy-to-understand explanations of difficult terms and context.
[0926] "Speech conversion means" refers to a technology that converts the generated commentary and information into audio data and provides it to the reader.
[0927] The present invention is a glasses-type device system designed to enhance the reading experience. The device combines a front-mounted image sensor, an eye tracking device, an artificial intelligence means, a background music selection means, a bone conduction sound device, a commentary generation means, and a voice conversion means to provide optimal reading support to users. Detailed embodiments of the system are described below.
[0928] When a user puts on the AI reading glasses, the device automatically starts up. The device includes an eye tracker to track the user's gaze and a forward-facing image sensor to capture the text being read. An emotion engine then starts up, analyzing the user's facial expressions and tone of voice.
[0929] The device captures the page of the book the user is looking at with a front image sensor and sends the information to the server. At the same time, it acquires gaze information using an eye tracker and collects emotion data with an emotion engine. This information is continuously sent to the server.
[0930] The server converts the image data into text using OCR (optical character recognition) technology, such as software like Tesseract OCR. The acquired text data is combined with gaze data to identify the user's current reading position. At the same time, a sentiment analysis algorithm is used to analyze the emotion data and determine the user's current emotional state.
[0931] The server selects appropriate background music based on the analyzed content. For example, if you are feeling tense while reading, it will select relaxing music, and if you are enjoying yourself, it will select music that will enhance the atmosphere. The background music data is sent to the terminal and provided to the user via a bone conduction audio device.
[0932] Additionally, if the server detects words or contexts that are difficult to understand from the gaze data and text data, it generates a meaning or summary of those words or contexts. This content is generated using a generative AI model (such as the GPT series). The generated information is converted into audio data and sent from the server to the device. The device then provides the audio data to the user through a bone conduction audio device. If necessary, the information is also displayed on a visual feedback panel.
[0933] As a concrete example, consider a case where a user is reading a mystery fiction book. As the climax approaches, the user's heart rate increases. The device captures this scene through the forward-facing image sensor and acquires gaze information and emotional information.
[0934] The server receives the captured text, gaze information, and emotion data and analyzes it to determine whether it is a climax scene.
[0935] The server selects relaxing music to relieve tension based on the user's emotional state and sends it to the terminal.
[0936] The terminal provides the music to the user through a bone conduction audio device.
[0937] If a mysterious character called "Detective Albion" appears in the climax scene,
[0938] The server determines that this character is difficult for the user to understand, and uses a generative AI model to generate meaning and background, such as "Detective Albaio is a mysterious detective and a character with a hidden past."
[0939] The terminal provides this information to the user through a bone conduction audio device.
[0940] This system allows users to receive appropriate music and explanations while reading, allowing them to become more immersed in the world of the book. In addition, by linking it with a generative AI model, it is possible to generate prompts to aid the reader in comprehension.
[0941] Prompt Sentence Examples
[0942] Please provide a brief background on the character "Detective Albio."
[0943] Choose relaxing music to listen to while you read.
[0944] The above is an embodiment of the present invention, which allows users to have a more comfortable and immersive reading experience.
[0945] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0946] Step 1:
[0947] The user puts on the AI reading glasses. The device automatically starts up, activating the eye tracker to track the user's gaze and the forward-facing image sensor to capture the text being read. The emotion engine also starts up.
[0948] Specifically, the device starts up the image sensor and eye tracker, and goes through an initialization process to prepare for acquiring gaze and video data.
[0949] Input: The user puts on the device.
[0950] Output: Eye tracking and image acquisition ready.
[0951] Step 2:
[0952] The device uses a forward-facing image sensor to capture the page of the book the user is reading, an eye tracker to collect gaze information in real time, and an emotion engine to extract emotional data from facial expressions and tone of voice.
[0953] Specifically, the device continuously captures images of the page using a front image sensor, acquires user gaze data using an eye tracker, and performs facial expression and voice analysis using an emotion engine.
[0954] Input: User's gaze information, forward image, emotion data.
[0955] Output: Captured image data, gaze data, emotion data.
[0956] Step 3:
[0957] The device continuously transmits captured image data, gaze data, and emotion data to the server.
[0958] Specifically, the terminal compresses the collected data at regular intervals and transmits it to the server via packet communication. Since the data is transmitted in real time, network bandwidth and communication efficiency are taken into consideration.
[0959] Input: Captured image data, gaze data, emotion data.
[0960] Output: Various data sent to the server.
[0961] Step 4:
[0962] The server converts the received image data into text using OCR technology (e.g., Tesseract OCR), combines it with gaze data to identify the part the user is reading, and analyzes emotion data to determine the user's emotional state.
[0963] Specifically, the server converts image data into text using an OCR engine, maps it with gaze data, and then uses an emotion analysis algorithm to identify the user's emotional state.
[0964] Input: Transmitted image data, gaze data, emotion data.
[0965] Output: Transformed text data, identified reading passages, and the user's emotional state.
[0966] Step 5:
[0967] The server selects appropriate background music based on the analysis results. For example, if the user is feeling tense, it selects relaxing music, and if the user is having fun, it selects music that will liven up the atmosphere. The selected background music data is sent to the device.
[0968] Specifically, the server selects the most suitable music from multiple music databases, compresses the music data, and sends it to the terminal.
[0969] Input: Parsed reading content, user's emotional state.
[0970] Output: Selected background music data, sent to the device.
[0971] Step 6:
[0972] The background music data received by the terminal is played through a bone conduction audio device.
[0973] Specifically, the device stores background music data in a buffer and plays it in real time through a bone conduction audio device.
[0974] Input: BGM data sent from the server.
[0975] Output: Providing background music to the user.
[0976] Step 7:
[0977] The server uses gaze data and text data to identify words and contexts that the user has difficulty comprehending, and generates meanings and summaries of those words and contexts. The generated information is converted into audio data and sent to the device.
[0978] Specifically, the server analyzes gaze data and text data, generates meaning and summaries using a generative AI model (e.g., GPT-3), and converts this text data into audio data.
[0979] Input: gaze data, text data.
[0980] Output: Generated description data and audio data.
[0981] Step 8:
[0982] The audio data received by the terminal is provided to the user through a bone conduction audio device.
[0983] Specifically, the device decodes the audio data, plays the audio through a bone conduction audio device, and optionally displays visual information on the feedback panel.
[0984] Input: Audio data sent from the server.
[0985] Output: Providing audio information to the user.
[0986] In this way, through the specific processing performed at each step, the user can receive appropriate reading assistance in real time.
[0987] (Application example 2)
[0988] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0989] Conventional reading assistance devices and electronic payment systems are unable to reflect the user's emotional state or gaze data in real time, limiting their ability to provide diverse information and improve the user experience. Furthermore, when shopping online, users often feel stressed when choosing the most suitable product and payment method, which can reduce their motivation to purchase. There is a need to solve these issues and provide users with a richer experience and more efficient shopping support.
[0990] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0991] In this invention, the server includes: means for detecting the user's gaze direction using a camera installed in front; gaze measurement means for collecting user gaze information in real time; artificial intelligence means for analyzing the content of reading and online activities based on the collected gaze information and data from the front camera; background music selection means for selecting appropriate background music based on the analyzed content; bone conduction speaker means for playing the selected background music; explanation generation means for generating and providing explanations of difficult words and contexts as necessary; product recommendation means for capturing the product page the user is viewing while shopping online and recommending appropriate products and services based on that information and emotional data; and payment assistance means for suggesting the optimal payment method based on the analyzed emotional data and supporting purchasing behavior. This enables optimal music selection and product recommendation based on emotional analysis.
[0992] A "front camera" is a photographic device installed on the front of a device to capture the user's line of sight and what they are viewing.
[0993] "Gaze measurement means" refers to technology or equipment that collects a user's gaze information in real time and analyzes its direction and focus.
[0994] "Artificial intelligence means" refers to technologies and systems that analyze collected gaze information and data from the front camera to understand the content of reading and online activities.
[0995] The "background music selection means" refers to a technology or system for selecting optimal background music based on the analyzed content.
[0996] "Bone conduction speaker means" means an audio device that uses bone conduction technology to transmit selected background music to a user.
[0997] "Explanation generation means" refers to technology or systems that generate explanations of difficult words or contexts as needed and provide them to users.
[0998] A "product recommendation method" is a technology or system that captures the product page a user is viewing while shopping online and recommends appropriate products or services based on that information and emotional data.
[0999] "Payment assistance methods" are technologies and systems that suggest the most suitable payment method to users based on analyzed emotional data and assist them in their purchasing behavior.
[1000] The present invention is implemented as a system for improving the reading experience and online shopping experience by a device that combines a forward-facing camera, a gaze measurement means, an artificial intelligence means, a background music selection means, a bone conduction speaker means, a description generation means, a product recommendation means, and a payment assistance means.
[1001] 1. Hardware configuration:
[1002] The front-facing camera is located at the front of the device to capture the user's line of sight and what they are viewing.
[1003] The gaze measurement means is used to collect user gaze information in real time.
[1004] The bone conduction speaker is used to convey selected background music and explanatory information to the user.
[1005] 2. Software configuration:
[1006] The artificial intelligence means analyzes the collected gaze information and data from the front camera to understand the content of your reading and online activities.
[1007] The background music selection means selects appropriate background music based on the analyzed content.
[1008] The explanation generating means generates explanations of difficult words and contexts as needed and provides them to the user.
[1009] The product recommendation means captures the product page that the user is viewing while shopping online and recommends appropriate products and services based on that information and emotional data.
[1010] The payment assistance tool suggests the most suitable payment method to the user based on the analyzed emotional data, and supports purchasing behavior.
[1011] 3. Data processing and calculation:
[1012] The video data acquired from the front camera is sent to a server and converted into text data using OCR technology.
[1013] The gaze data and emotion analysis data acquired by the gaze measurement means are analyzed by the artificial intelligence means to identify the specific part the user is reading.
[1014] The background music selection means selects the most suitable music based on the emotion and theme of the analyzed text and plays it through a bone conduction speaker.
[1015] The product recommendation and payment assistance methods capture the product page the user is viewing while shopping online, and suggest the most suitable products and payment methods based on that information and emotional data.
[1016] 4. Specific examples:
[1017] For example, if a user is reading a mystery novel, emotion analysis can tell if the user is feeling tense. The system can detect this and play relaxing music through a bone conduction speaker. It can also generate and provide audio explanations of difficult words and context that appear in the climax scene.
[1018] 5. Example of a generative AI model and prompt:
[1019] Example prompt sentence:
[1020] Capture the product page the user is viewing with a camera, collect gaze and emotion data, and send that data to your server to recommend the best products and payment methods.
[1021] This allows users to receive optimal music and descriptions, product recommendations, and payment assistance while reading or shopping online, greatly improving their experience.
[1022] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1023] Step 1:
[1024] The user puts on the device ("Smart Payment Assist Glasses") and the system starts automatically. The input is the user's wearing information. The device activates the camera and gaze measurement means and begins to capture the user's gaze information and facial expressions in real time. The output obtained is gaze data and facial expression data captured in real time.
[1025] Step 2:
[1026] The device uses the front camera to capture an image of the product page the user is viewing. The input is the product page the user is viewing. The image data captured by the camera is sent to the server. The output is the image data sent to the server.
[1027] Step 3:
[1028] The server analyzes the received image data using OCR (Optical Character Recognition) technology to extract text data related to the product. The input is image data captured by a camera. Through OCR analysis, text data is extracted to obtain the specific product information the user is viewing. The output is the extracted text data.
[1029] Step 4:
[1030] The device uses gaze measurement means to collect the user's gaze information in real time. The input is the user's eye movement and focus. The gaze data is continuously sent to a server and used to identify the specific location the user is looking at. The output is the identified gaze information.
[1031] Step 5:
[1032] The server uses an emotion engine to analyze the user's facial expressions and tone of voice to obtain current emotion data. The inputs are the user's facial expressions and tone of voice. The emotion data is analyzed along with the gaze data to help identify the user's current emotional state. The output is the identified emotion data.
[1033] Step 6:
[1034] The server comprehensively analyzes gaze data, text data, and emotion data to determine the products in which the user is interested and the degree of their willingness to purchase them. The inputs are gaze data, text data, and emotion data. The analysis results in the recommendation of appropriate products and services. The output is information about recommended products and services.
[1035] Step 7:
[1036] Based on the analyzed content, the server selects appropriate background music using a background music selection means. The inputs are the user's emotional state and the text content being viewed. The selected music is sent to the terminal and played to the user through a bone conduction speaker. The resulting output is the selected background music.
[1037] Step 8:
[1038] The server generates explanations of difficult words and contexts for the user and converts them into audio data. The input is the text content the user is viewing and the identified difficult parts. The explanation generation means generates a summary and sends it to the terminal as audio data. The output obtained is explanatory information as audio data.
[1039] Step 9:
[1040] The server proposes the optimal payment method based on the analyzed emotional data and purchasing intention information and notifies the user. The input is the user's emotional state and recommended product information. The payment assistance means selects the appropriate payment method and provides assistance to the user. The output obtained is the proposed payment method.
[1041] This allows users to receive the best music and explanations, product recommendations, and payment assistance while reading or shopping online.
[1042] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1043] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1044] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1045] [Fourth embodiment]
[1046] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1047] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1048] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1049] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1050] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1051] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1052] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1053] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1054] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1055] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1056] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1057] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1058] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1059] The present invention provides a glasses-type reading assistance device for improving the reading experience. The device includes a front camera, a gaze measurement unit, an artificial intelligence unit, a background music selection unit, a bone conduction speaker, and a description generation unit. The program processing of this system is explained below in natural language.
[1060] 1. Setup Phase
[1061] The user puts on the "AI reading glasses." The device automatically starts up and activates the eye tracker to follow the user's gaze and the front camera to capture the text being read. The device is then ready for reading.
[1062] 2. Real-time understanding of reading status
[1063] The device captures the page of the book the user is looking at with the front camera and sends the information to the server. At the same time, the device continues to acquire gaze information using the eye tracker, and this information is continuously sent to the server.
[1064] 3. Analysis of reading content
[1065] The server receives the text and gaze data and analyzes it. This analysis identifies which part the user is reading and understands what they are reading. Based on this information, the server can select appropriate background music or detect difficult passages.
[1066] 4. Select and play appropriate background music
[1067] The server selects appropriate background music based on the reading content analyzed. The selected background music data is sent to the device, which then uses a bone conduction speaker to provide the background music to the user. This allows music to be played according to the situation, enhancing the reading atmosphere.
[1068] 5. Assistance with difficult-to-read passages
[1069] If the server determines that a particular word or context is difficult for the user to understand, it generates a meaning or summary of that word or context, sends this information to the device, and provides it to the user audibly through a bone conduction speaker. If necessary, it may also be displayed as visual feedback on the feedback panel.
[1070] Specific operation example
[1071] Reading situations
[1072] A user is reading a historical novel and is about to enter a battle scene. The device captures the scene through the front camera and acquires gaze information.
[1073] The server receives the captured text and gaze information and analyzes it to determine that it is a battle scene.
[1074] The server selects epic battle music that heightens the tension and sends it to the device.
[1075] The device provides the music to the user through a bone conduction speaker.
[1076] If the difficult word "Yamato Taro Yoshiie" appears in the middle of a battle scene,
[1077] The server determines that this word is difficult for the user to understand and generates a meaning and background such as "Hachiman Taro Yoshiie was a military commander in the Heian period, and was later enshrined as a god at Hachiman Shrine."
[1078] The device provides this information to the user through a bone conduction speaker.
[1079] This allows users to immerse themselves even more deeply in the world of the book through music and explanations.
[1080] As can be seen, the present invention enhances the reading experience and allows readers to become immersed in the world of the book.
[1081] The processing flow will be explained below.
[1082] Step 1:
[1083] The user puts on the "AI reading glasses." The device automatically starts up and activates the eye tracker to follow the user's gaze and the front camera to capture the text being read.
[1084] Step 2:
[1085] The device uses the front camera to capture the page of the book the user is looking at, formats the captured image data, and saves it in temporary storage.
[1086] Step 3:
[1087] The device uses an eye tracker to acquire user gaze information and temporarily stores the gaze data.
[1088] Step 4:
[1089] The device sends the captured image data and gaze data as a set to the server in real time.
[1090] Step 5:
[1091] The server converts the received image data into text data through an OCR (optical character recognition) process. Based on the converted text data and gaze data, the server performs natural language analysis to identify what the user is reading and where they are reading from.
[1092] Step 6:
[1093] Based on the analyzed text data, the server selects the appropriate background music for the scene. For example, if it's a battle scene, it will select epic battle music.
[1094] Step 7:
[1095] The server sends the selected background music data to the device, which then plays it through the bone conduction speaker, allowing users to enjoy music suitable for reading.
[1096] Step 8:
[1097] The server uses gaze data and text data to detect words and contexts that are difficult to understand, such as difficult historical or academic terms.
[1098] Step 9:
[1099] The server generates meaning and context for each difficult word and context that is detected, which is then converted into audio data.
[1100] Step 10:
[1101] The server sends the generated audio data to the device, which then plays it back through a bone conduction speaker and optionally displays visual information on a feedback panel.
[1102] Through the above steps, the present invention improves the reading experience and provides an environment where users can immerse themselves in the world of books.
[1103] Example 1
[1104] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1105] The modern reading experience remains static, making it difficult for readers to deeply understand the meaning of text or customize their reading environment. Furthermore, there is a lack of ways to quickly understand the meaning of difficult words or contexts when encountering them. Furthermore, there is no function to select appropriate background music, making it difficult to create an environment conducive to immersion in reading.
[1106] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1107] In this invention, the server includes means for detecting the user's gaze direction using a camera installed in front, gaze tracking means for collecting the user's gaze information in real time, machine learning means for analyzing the content being read based on the collected gaze information and data from the camera, background sound selection means for selecting appropriate background sound based on the analyzed content, bone conduction sound means for playing the selected background sound, and explanation generation means for generating and providing explanations of difficult words and contexts as needed, thereby improving the reading experience, enabling the reader to deeply understand the content of the book, and providing an appropriate acoustic environment.
[1108] The "photography device" is a device installed in front of the user to detect the direction of the user's line of sight.
[1109] "Eye tracking means" is a technology for collecting user eye gaze information in real time.
[1110] "Machine learning means" is a technology that analyzes the content of reading based on collected gaze information and data from a camera.
[1111] The "background sound selection means" is a technique for selecting appropriate background sound based on the analyzed content.
[1112] "Bone conduction sound means" is a technology for reproducing selected background sounds.
[1113] "Explanation generator" is a technology for generating and providing explanations of difficult words and contexts as needed.
[1114] The present invention relates to a glasses-type reading assistance device for improving reading experience, which includes a front-mounted camera, an eye tracking unit, a machine learning unit, a background sound selection unit, a bone conduction sound unit, and an explanation generation unit.
[1115] When a user puts on the eyeglass-type reading assistance device, the device automatically starts up and activates the eye tracker (eye tracking means) to track the user's gaze and the front camera (photography device) to capture the text being read, making the device ready for reading.
[1116] Next, the device captures the page of the book the user is looking at with the front camera and sends the information to the server. At the same time, the device continues to acquire gaze information using the eye tracker, which is then continuously sent to the server.
[1117] The server analyzes the transmitted gaze data and captured text images. It uses machine learning methods to identify which part the user is reading and understand what is being read. Based on this information, the server can select appropriate background audio or detect difficult words or contexts.
[1118] The server then selects appropriate background sounds based on the analysis results. The selected background sound data is sent to the device, which then provides the sound to the user using bone conduction audio, thereby playing music appropriate to the situation and enhancing the reading atmosphere.
[1119] Furthermore, if the server determines that a particular word or sentence is difficult for the user to understand as a result of the analysis, it generates a meaning or summary of the word or sentence. This information is sent to the terminal, which then provides it to the user audibly through bone conduction acoustic means. If necessary, the terminal may also display the information on a feedback panel as visual feedback.
[1120] Specific examples
[1121] For example, suppose a user is reading a historical novel. When the user comes to a battle scene in the book, the device captures the scene through the front camera and acquires gaze information. The server receives the captured text and gaze information and analyzes that it is a battle scene. As a result, the server selects epic battle sounds to heighten the tension and transmits them to the device. The device then provides this sound to the user through bone conduction audio.
[1122] Also, if the user comes across a difficult word like "Yoshiie Hachiman" while reading, the server will determine that the word is difficult for the user to understand and generate a meaning and background such as "Yoshiie Hachiman was a military commander in the Heian period, and was later enshrined as a god at Hachiman Shrine." The terminal will provide this information to the user through bone conduction acoustic means.
[1123] Prompt Sentence Examples
[1124] "Please select background music that matches the battle scenes in a historical novel and generate commentary for Yawata Taro Yoshiie."
[1125] As described above, the present invention improves the reading experience, allows the reader to understand the contents of the book more deeply, and provides an appropriate acoustic environment.
[1126] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1127] Step 1:
[1128] The user puts on the "glasses-type reading assistance device." The device automatically starts up, and the eye tracker (a means of tracking gaze) and the front camera (a photographing device) for capturing the text being read are enabled. This allows the device to obtain gaze information in real time, preparing the user's reading environment.
[1129] Step 2:
[1130] The device captures the page of the book the user is looking at through the front camera. At the same time, the eye tracker collects the user's gaze information. As input, the device receives image data and gaze data from the front camera, and sends these data to the server in real time. As output, the captured image data and gaze data are sent to the server.
[1131] Step 3:
[1132] The server analyzes the received gaze data and captured image. It uses gaze data and captured image as input. It uses machine learning algorithms to identify which part the user is reading and performs text analysis. The output is the specific information about the reading part and the analysis results.
[1133] Step 4:
[1134] The server selects appropriate background audio based on the analysis results. It uses the specific information about the reading section and the analysis results as input. The server uses a background audio selection algorithm to select music that matches the analyzed content. The selected background audio data is generated as output.
[1135] Step 5:
[1136] The server sends the selected background sound data to the terminal. The background sound data is used as input. The terminal receives this sound data and provides music to the user using bone conduction sound means. As output, appropriate background music is played for the user.
[1137] Step 6:
[1138] The server identifies difficult words and sentences and generates their meaning and context. It uses gaze data and text analysis results as input. It uses an explanation generation algorithm to create explanations for difficult words and contexts. The output is the generated explanation data.
[1139] Step 7:
[1140] The server sends the generated explanation data to the terminal, which uses the explanation data as input. The terminal receives the explanation data and provides audio explanations to the user through bone conduction acoustic means, and also provides visual feedback if necessary. As output, the user is provided with audio and visual explanations of difficult words and contexts.
[1141] Through the above processing steps, the present invention can improve the reading experience, allow users to understand the contents of the book more deeply, and provide a suitable acoustic environment.
[1142] (Application example 1)
[1143] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1144] Traditional reading experiences have the drawback of making it difficult for readers to gain a deep understanding of the content or experience a sense of realism. In particular, when encountering difficult words or context, there are limited ways to receive on-the-spot explanations, which frequently interrupts the flow of reading. There is also a demand for incorporating background music and sound effects into reading to improve the quality of the reading experience.
[1145] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1146] In this invention, the server includes: means for detecting the reader's gaze direction using a camera installed in front; gaze measurement means for collecting the reader's gaze information in real time; artificial intelligence means for analyzing the content being read based on the collected gaze information and data from the front camera; background music selection means for selecting appropriate background music based on the analyzed content; a sound wave conduction audio device for playing the selected background music; explanation generation means for generating and providing explanations of difficult words and context as needed; data communication means for transmitting the text and gaze information being read to the data server, the server analyzing the transmitted data, and providing the user with the analysis results; and audio control means for transmitting the background music selected based on the analysis results to the smart device and for the smart device to play the background music. This significantly improves the quality of the reading experience, allowing the reader to deeply understand the content and experience a sense of realism.
[1147] 1. "Camera" refers to a device that is installed in front of the user to detect the direction of their gaze and capture images in real time.
[1148] 2. "Gaze measurement means" refers to technology or equipment used to collect readers' gaze information in real time and analyze that information.
[1149] 3. "Artificial intelligence means" refers to algorithms and processes used to analyze the content of a reading based on collected gaze information and camera data.
[1150] 4. "Background music selection means" refers to a method or system for selecting appropriate background music based on the analyzed reading content.
[1151] 5. "Sound wave conduction sound device" means a bone conduction speaker or similar sound reproduction device that provides a user with selected background music.
[1152] 6. "Explanation generation means" refers to technology that generates explanations for difficult words or contexts and provides them to users in audio or visual form.
[1153] 7. "Data communication means" refers to the network technology and protocols used to transmit text and eye gaze information during reading to a data server and receive analysis results.
[1154] 8. "Audio Control Means" means the control technology or algorithm for playing background music selected based on the analysis results on a smart device.
[1155] 9. "Data server" refers to a computer system installed on the cloud or elsewhere that analyzes transmitted data and provides the results.
[1156] To implement this invention, the following system configuration and its operating procedures will be specifically described. The system includes a camera, a gaze measurement means, an artificial intelligence means, a background music selection means, a sound-transmitting audio device, a description generation means, a data communication means, and an audio control means.
[1157] Hardware and Software Use
[1158] Camera: Uses the front-facing camera on your smartphone or tablet to capture where you're looking and the text you're reading.
[1159] Eye tracking method: Use software with eye tracking functionality (e.g., the GazeTracking library).
[1160] Artificial intelligence measures: Machine learning models are used to analyze collected gaze information and captured text data.
[1161] Background music selection method: Use a music library and music recommendation algorithms to select background music appropriate for the reading content.
[1162] Sound wave conduction sound device: Plays background music using bone conduction speakers or smartphone speakers.
[1163] Explanation generation: Use text generation models and text-to-speech technologies (e.g., Google Cloud Text-to-Speech API) to generate explanations for difficult words and contexts and convert them into audio.
[1164] Data communication means: An internet communication protocol (e.g., HTTP request) is used to send text data and gaze data to the cloud server and return the analysis results to the smart device.
[1165] Sound control means: Use the music playback library to play background music on the smart device based on the analysis results sent from the cloud server.
[1166] Specific examples
[1167] Reading situations
[1168] Assume a user is reading a historical novel on their smartphone. When they reach a battle scene in the book, the system works as follows:
[1169] 1. Eye Tracking and Text Capture:
[1170] The page the user is reading is captured through the smartphone camera.
[1171] At the same time, the gaze measurement means collects the user's gaze data in real time.
[1172] 2. Data analysis and background music selection:
[1173] The captured text data and gaze data are transmitted to a cloud server.
[1174] The cloud server uses machine learning models to analyze the data and identify that the user is reading a fight scene.
[1175] Based on the analysis results, epic battle music is selected and the music data is sent back to the smartphone.
[1176] 3. Play background music and explain difficult passages:
[1177] The smartphone plays battle music through a sonic wave transmission audio device.
[1178] If the difficult word "Yahata Taro Yoshiie" appears in the middle of a battle scene, the cloud server will determine that the word is difficult to understand and generate its meaning and background.
[1179] The smartphone plays the generated explanation as audio and provides it to the user.
[1180] Prompt Sentence Examples
[1181] A user is reading a historical novel and is currently approaching a battle scene. Choose epic battle music to match this scene and explain the meaning and background of the difficult passage "Yaman Taro Yoshiie."
[1182] This will significantly improve the quality of the reading experience, allowing readers to gain a deeper understanding of the content and experience a more immersive experience.
[1183] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1184] Step 1:
[1185] The user opens the smartphone, launches the app, and selects an e-book. The input is the e-book file selected by the user and the initial gaze information data. The output is a state where camera capture and gaze tracking are ready. The camera is placed in front and the eye tracking function is enabled.
[1186] Step 2:
[1187] The device uses a front camera to capture text data of the page the user is reading and collects gaze data in real time through an eye gaze measurement means. The input is the camera image and real-time gaze position data. The output is the captured text data and gaze data.
[1188] Step 3:
[1189] The captured text data and gaze data are sent from the terminal to the server. The input is the text data and gaze data sent from the terminal. The output is a notification that the text data and gaze data have been received. These data are sent to the cloud server using a data communication means.
[1190] Step 4:
[1191] The server analyzes the received text data and gaze data using artificial intelligence. The input is the text data and gaze data. The output is the analysis result of what the user is reading. For example, it may be determined that the user is reading a battle scene.
[1192] Step 5:
[1193] The server selects appropriate background music based on the analysis results. The input is the analyzed reading content (e.g., a battle scene). The output is a URL or file of the selected background music data. Using the background music selection means, epic music suitable for the battle scene is selected.
[1194] Step 6:
[1195] The server sends the selected background music data to the terminal. The input is the URL or file of the background music data. The output is the background music data sent to the terminal. The music data is sent to the smartphone using a data communication means.
[1196] Step 7:
[1197] The background music data received by the terminal is played on an acoustic wave conduction acoustic device. The input is the received background music data. The output is the background music being listened to by the user. Using the acoustic control means, the background music is played on the smartphone's bone conduction speaker.
[1198] Step 8:
[1199] The server identifies difficult-to-read passages and generates their meaning and background using an explanation generation means. The input is the analyzed text data and information on the identified difficult-to-read passages. The output is the generated explanatory text. For example, the text generated for "Hachiman Taro Yoshiie" is "He was a military commander in the Heian period, and later enshrined as a god at Hachiman Shrine."
[1200] Step 9:
[1201] The server transmits the explanation generated by the explanation generation means to the terminal. The input is the generated explanation text. The output is the explanation audio data transmitted to the terminal. The explanation data is transmitted to the smartphone using the data communication means.
[1202] Step 10:
[1203] The terminal plays the received explanatory data on an acoustic wave conduction audio device. The input is the received explanatory audio data. The output is the explanatory audio that the user is listening to. Using the audio control means, the explanation is played through a bone conduction speaker.
[1204] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1205] This invention is a glasses-type reading assistance device that combines a front camera, gaze tracking means, artificial intelligence means, background music selection means, bone conduction speakers, explanation generation means, and an emotion engine to improve the reading experience. This device analyzes the content being read based on the user's gaze and emotions, plays background music appropriate to the situation, and provides support for difficult words and context. The program processing of this system is explained below in natural language.
[1206] 1. Setup Phase
[1207] The user puts on the AI reading glasses. The device automatically activates the eye tracker to track the user's gaze and the front camera to capture the text being read. It also activates an emotion engine that analyzes the user's facial expressions and tone of voice.
[1208] 2. Real-time understanding of reading status
[1209] The device captures the page of the book the user is looking at with the front camera and sends the information to the server. At the same time, the device continues to acquire gaze information using the eye tracker. The emotion engine also analyzes the user's facial expressions and tone of voice to acquire emotional data. This information is continuously sent to the server.
[1210] 3. Analysis of reading content
[1211] The server receives the text, gaze data, and emotion data and analyzes them. The server recognizes the text data using OCR (optical character recognition) technology and combines it with the gaze data to identify the specific part the user is reading. It also analyzes the emotion data to determine the user's current emotional state.
[1212] 4. Select and play appropriate background music
[1213] The server selects appropriate background music based on the reading content and the user's emotional data analyzed. For example, if the user is feeling tense while reading, it selects relaxing music, and if the user is enjoying themselves, it selects music that enhances the mood. The selected background music data is sent to the device, which then uses a bone conduction speaker to provide the background music to the user.
[1214] 5. Assistance with difficult-to-read passages
[1215] If the server detects words or contexts that are difficult to read from the gaze data and text data, it generates a meaning and summary of the words or contexts. The generated information is converted into audio data and sent from the server to the device. The device then provides the audio data to the user through a bone conduction speaker. It may also display visual information on a feedback panel as needed.
[1216] Specific operation example
[1217] Reading situations
[1218] A user is reading a mystery novel. The climax of the story is approaching, and the user's heart rate is rising. The device captures the scene through the front camera and acquires gaze and emotion information.
[1219] The server receives the captured text, gaze information, and emotion data and analyzes it to determine whether it is a climax scene.
[1220] The server selects relaxing music to relieve tension based on the user's emotional state and sends it to the terminal.
[1221] The device provides the music to the user through a bone conduction speaker.
[1222] When the difficult name "Albaiohiko" appears in the climax scene,
[1223] The server determines that this name is difficult for the user to understand and generates a meaning and background, such as "Albaiohiko is a mysterious detective and a character with a hidden past."
[1224] The device provides this information to the user through a bone conduction speaker.
[1225] This allows users to receive appropriate music and explanations while reading, allowing them to become more immersed in the world of the book.
[1226] As described above, the present invention improves the reading experience and provides an environment in which users can immerse themselves in the world of books.
[1227] The processing flow will be explained below.
[1228] Step 1:
[1229] The user puts on the "AI reading glasses" and turns on the device, which activates the eye tracker, front-facing camera, and emotion engine.
[1230] Step 2:
[1231] The device uses the front camera to periodically capture the page of the book the user is looking at, and the captured image data is temporarily stored.
[1232] Step 3:
[1233] The device uses an eye tracker to acquire the user's gaze information in real time and temporarily stores this gaze data.
[1234] Step 4:
[1235] The device uses an emotion engine to analyze the user's facial expressions and tone of voice to determine their current emotional state, and emotion data is also stored.
[1236] Step 5:
[1237] The terminal transmits the captured image data, gaze data, and emotion data as a set to the server in real time.
[1238] Step 6:
[1239] The server converts the received image data into text data using OCR (optical character recognition) technology, and then analyzes the text data by combining it with gaze data and emotion data.
[1240] Step 7:
[1241] Based on the analysis, the server identifies the specific part the user is reading and makes a comprehensive assessment of the user's emotional state.
[1242] Step 8:
[1243] The server selects appropriate background music based on the analysis results and emotional data. For example, it selects relaxing music for tense scenes and music that enhances the atmosphere for happy scenes.
[1244] Step 9:
[1245] The server sends the selected background music data to the device, which then uses a bone conduction speaker to provide the music to the user.
[1246] Step 10:
[1247] The server detects difficult words and contexts from gaze data and text data, and generates meaning and background information for the detected content.
[1248] Step 11:
[1249] The server converts the generated explanation data into audio data and sends it to the terminal, which then provides the audio data to the user through a bone conduction speaker. The terminal also displays visual information on a feedback panel as needed.
[1250] As a specific example of how it works, when a user is reading a mystery novel and reaches a tense climax, the emotion engine detects the user's tension.
[1251] The server selects relaxing background music and sends it to the device.
[1252] The device plays the background music through a bone conduction speaker, helping the user relax.
[1253] When the difficult name "Albaiohiko" appears, the server generates background information about the name and sends it to the terminal as voice data.
[1254] The device plays audio to help the user understand.
[1255] In this way, the present invention takes into account the user's emotions while providing appropriate music and information support to enhance the reading experience.
[1256] Example 2
[1257] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1258] While various technologies have been proposed to improve the reading experience, systems that provide real-time support based on changes in the reader's gaze and emotions are not yet fully developed. In particular, technologies that provide appropriate background music while reading and clearly explain difficult words and contexts are in need of further development. This creates a demand for systems that can enhance immersion and comprehension while reading.
[1259] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1260] In this invention, the server includes means for detecting the reader's gaze direction using an image sensor installed in front, a gaze measurement device for collecting the reader's gaze information in real time, artificial intelligence means for analyzing the content being read based on the collected gaze information and data from the front image sensor, background music selection means for selecting appropriate background music based on the analyzed content and the reader's emotional data, a bone conduction audio device for playing the selected background music, explanation generation means for generating and providing explanations of difficult terms and context as necessary, and voice conversion means for converting the analyzed data into voice data and providing it to the reader. This makes it possible to provide dynamic support in real time according to the reader's gaze and emotions, significantly improving the quality of the reading experience.
[1261] A "forward-mounted image sensor" is an image capture device mounted in front of the reader to detect the reader's line of sight.
[1262] The "means for detecting the reader's line of sight" is a mechanism that uses an image sensor installed in front to identify where the reader is looking.
[1263] An "eye gaze measurement device that collects reader gaze information in real time" is a device that instantly tracks the reader's eye movements and collects that data.
[1264] "Artificial intelligence means" refers to technology that analyzes collected gaze information and data from the forward image sensor to understand the content and behavior of the reader while reading.
[1265] The "background music selection means" is a mechanism for selecting optimal background music based on the analyzed content and emotional data.
[1266] A "bone conduction audio device" is an audio playback device that allows selected background music to be heard directly through the bones.
[1267] "Explanation generation means" is a technology that generates the meaning and summary of text to provide easy-to-understand explanations of difficult terms and context.
[1268] "Speech conversion means" refers to a technology that converts the generated commentary and information into audio data and provides it to the reader.
[1269] The present invention is a glasses-type device system designed to enhance the reading experience. The device combines a front-mounted image sensor, an eye tracking device, an artificial intelligence means, a background music selection means, a bone conduction sound device, a commentary generation means, and a voice conversion means to provide optimal reading support to users. Detailed embodiments of the system are described below.
[1270] When a user puts on the AI reading glasses, the device automatically starts up. The device includes an eye tracker to track the user's gaze and a forward-facing image sensor to capture the text being read. An emotion engine then starts up, analyzing the user's facial expressions and tone of voice.
[1271] The device captures the page of the book the user is looking at with a front image sensor and sends the information to the server. At the same time, it acquires gaze information using an eye tracker and collects emotion data with an emotion engine. This information is continuously sent to the server.
[1272] The server converts the image data into text using OCR (optical character recognition) technology, such as software like Tesseract OCR. The acquired text data is combined with gaze data to identify the user's current reading position. At the same time, a sentiment analysis algorithm is used to analyze the emotion data and determine the user's current emotional state.
[1273] The server selects appropriate background music based on the analyzed content. For example, if you are feeling tense while reading, it will select relaxing music, and if you are enjoying yourself, it will select music that will enhance the atmosphere. The background music data is sent to the terminal and provided to the user via a bone conduction audio device.
[1274] Additionally, if the server detects words or contexts that are difficult to understand from the gaze data and text data, it generates a meaning or summary of those words or contexts. This content is generated using a generative AI model (such as the GPT series). The generated information is converted into audio data and sent from the server to the device. The device then provides the audio data to the user through a bone conduction audio device. If necessary, the information is also displayed on a visual feedback panel.
[1275] As a concrete example, consider a case where a user is reading a mystery fiction book. As the climax approaches, the user's heart rate increases. The device captures this scene through the forward-facing image sensor and acquires gaze information and emotional information.
[1276] The server receives the captured text, gaze information, and emotion data and analyzes it to determine whether it is a climax scene.
[1277] The server selects relaxing music to relieve tension based on the user's emotional state and sends it to the terminal.
[1278] The terminal provides the music to the user through a bone conduction audio device.
[1279] If a mysterious character called "Detective Albion" appears in the climax scene,
[1280] The server determines that this character is difficult for the user to understand, and uses a generative AI model to generate meaning and background, such as "Detective Albaio is a mysterious detective and a character with a hidden past."
[1281] The terminal provides this information to the user through a bone conduction audio device.
[1282] This system allows users to receive appropriate music and explanations while reading, allowing them to become more immersed in the world of the book. In addition, by linking it with a generative AI model, it is possible to generate prompts to aid the reader in comprehension.
[1283] Prompt Sentence Examples
[1284] Please provide a brief background on the character "Detective Albio."
[1285] Choose relaxing music to listen to while you read.
[1286] The above is an embodiment of the present invention, which allows users to have a more comfortable and immersive reading experience.
[1287] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1288] Step 1:
[1289] The user puts on the AI reading glasses. The device automatically starts up, activating the eye tracker to track the user's gaze and the forward-facing image sensor to capture the text being read. The emotion engine also starts up.
[1290] Specifically, the device starts up the image sensor and eye tracker, and goes through an initialization process to prepare for acquiring gaze and video data.
[1291] Input: The user puts on the device.
[1292] Output: Eye tracking and image acquisition ready.
[1293] Step 2:
[1294] The device uses a forward-facing image sensor to capture the page of the book the user is reading, an eye tracker to collect gaze information in real time, and an emotion engine to extract emotional data from facial expressions and tone of voice.
[1295] Specifically, the device continuously captures images of the page using a front image sensor, acquires user gaze data using an eye tracker, and performs facial expression and voice analysis using an emotion engine.
[1296] Input: User's gaze information, forward image, emotion data.
[1297] Output: Captured image data, gaze data, emotion data.
[1298] Step 3:
[1299] The device continuously transmits captured image data, gaze data, and emotion data to the server.
[1300] Specifically, the terminal compresses the collected data at regular intervals and transmits it to the server via packet communication. Since the data is transmitted in real time, network bandwidth and communication efficiency are taken into consideration.
[1301] Input: Captured image data, gaze data, emotion data.
[1302] Output: Various data sent to the server.
[1303] Step 4:
[1304] The server converts the received image data into text using OCR technology (e.g., Tesseract OCR), combines it with gaze data to identify the part the user is reading, and analyzes emotion data to determine the user's emotional state.
[1305] Specifically, the server converts image data into text using an OCR engine, maps it with gaze data, and then uses an emotion analysis algorithm to identify the user's emotional state.
[1306] Input: Transmitted image data, gaze data, emotion data.
[1307] Output: Transformed text data, identified reading passages, and the user's emotional state.
[1308] Step 5:
[1309] The server selects appropriate background music based on the analysis results. For example, if the user is feeling tense, it selects relaxing music, and if the user is having fun, it selects music that will liven up the atmosphere. The selected background music data is sent to the device.
[1310] Specifically, the server selects the most suitable music from multiple music databases, compresses the music data, and sends it to the terminal.
[1311] Input: Parsed reading content, user's emotional state.
[1312] Output: Selected background music data, sent to the device.
[1313] Step 6:
[1314] The background music data received by the terminal is played through a bone conduction audio device.
[1315] Specifically, the device stores background music data in a buffer and plays it in real time through a bone conduction audio device.
[1316] Input: BGM data sent from the server.
[1317] Output: Providing background music to the user.
[1318] Step 7:
[1319] The server uses gaze data and text data to identify words and contexts that the user has difficulty comprehending, and generates meanings and summaries of those words and contexts. The generated information is converted into audio data and sent to the device.
[1320] Specifically, the server analyzes gaze data and text data, generates meaning and summaries using a generative AI model (e.g., GPT-3), and converts this text data into audio data.
[1321] Input: gaze data, text data.
[1322] Output: Generated description data and audio data.
[1323] Step 8:
[1324] The audio data received by the terminal is provided to the user through a bone conduction audio device.
[1325] Specifically, the device decodes the audio data, plays the audio through a bone conduction audio device, and optionally displays visual information on the feedback panel.
[1326] Input: Audio data sent from the server.
[1327] Output: Providing audio information to the user.
[1328] In this way, through the specific processing performed at each step, the user can receive appropriate reading assistance in real time.
[1329] (Application example 2)
[1330] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1331] Conventional reading assistance devices and electronic payment systems are unable to reflect the user's emotional state or gaze data in real time, limiting their ability to provide diverse information and improve the user experience. Furthermore, when shopping online, users often feel stressed when choosing the most suitable product and payment method, which can reduce their motivation to purchase. There is a need to solve these issues and provide users with a richer experience and more efficient shopping support.
[1332] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1333] In this invention, the server includes: means for detecting the user's gaze direction using a camera installed in front; gaze measurement means for collecting user gaze information in real time; artificial intelligence means for analyzing the content of reading and online activities based on the collected gaze information and data from the front camera; background music selection means for selecting appropriate background music based on the analyzed content; bone conduction speaker means for playing the selected background music; explanation generation means for generating and providing explanations of difficult words and contexts as necessary; product recommendation means for capturing the product page the user is viewing while shopping online and recommending appropriate products and services based on that information and emotional data; and payment assistance means for suggesting the optimal payment method based on the analyzed emotional data and supporting purchasing behavior. This enables optimal music selection and product recommendation based on emotional analysis.
[1334] A "front camera" is a photographic device installed on the front of a device to capture the user's line of sight and what they are viewing.
[1335] "Gaze measurement means" refers to technology or equipment that collects a user's gaze information in real time and analyzes its direction and focus.
[1336] "Artificial intelligence means" refers to technologies and systems that analyze collected gaze information and data from the front camera to understand the content of reading and online activities.
[1337] The "background music selection means" refers to a technology or system for selecting optimal background music based on the analyzed content.
[1338] "Bone conduction speaker means" means an audio device that uses bone conduction technology to transmit selected background music to a user.
[1339] "Explanation generation means" refers to technology or systems that generate explanations of difficult words or contexts as needed and provide them to users.
[1340] A "product recommendation method" is a technology or system that captures the product page a user is viewing while shopping online and recommends appropriate products or services based on that information and emotional data.
[1341] "Payment assistance methods" are technologies and systems that suggest the most suitable payment method to users based on analyzed emotional data and assist them in their purchasing behavior.
[1342] The present invention is implemented as a system for improving the reading experience and online shopping experience by a device that combines a forward-facing camera, a gaze measurement means, an artificial intelligence means, a background music selection means, a bone conduction speaker means, a description generation means, a product recommendation means, and a payment assistance means.
[1343] 1. Hardware configuration:
[1344] The front-facing camera is located at the front of the device to capture the user's line of sight and what they are viewing.
[1345] The gaze measurement means is used to collect user gaze information in real time.
[1346] The bone conduction speaker is used to convey selected background music and explanatory information to the user.
[1347] 2. Software configuration:
[1348] The artificial intelligence means analyzes the collected gaze information and data from the front camera to understand the content of your reading and online activities.
[1349] The background music selection means selects appropriate background music based on the analyzed content.
[1350] The explanation generating means generates explanations of difficult words and contexts as needed and provides them to the user.
[1351] The product recommendation means captures the product page that the user is viewing while shopping online and recommends appropriate products and services based on that information and emotional data.
[1352] The payment assistance tool suggests the most suitable payment method to the user based on the analyzed emotional data, and supports purchasing behavior.
[1353] 3. Data processing and calculation:
[1354] The video data acquired from the front camera is sent to a server and converted into text data using OCR technology.
[1355] The gaze data and emotion analysis data acquired by the gaze measurement means are analyzed by the artificial intelligence means to identify the specific part the user is reading.
[1356] The background music selection means selects the most suitable music based on the emotion and theme of the analyzed text and plays it through a bone conduction speaker.
[1357] The product recommendation and payment assistance methods capture the product page the user is viewing while shopping online, and suggest the most suitable products and payment methods based on that information and emotional data.
[1358] 4. Specific examples:
[1359] For example, if a user is reading a mystery novel, emotion analysis can tell if the user is feeling tense. The system can detect this and play relaxing music through a bone conduction speaker. It can also generate and provide audio explanations of difficult words and context that appear in the climax scene.
[1360] 5. Example of a generative AI model and prompt:
[1361] Example prompt sentence:
[1362] Capture the product page the user is viewing with a camera, collect gaze and emotion data, and send that data to your server to recommend the best products and payment methods.
[1363] This allows users to receive optimal music and descriptions, product recommendations, and payment assistance while reading or shopping online, greatly improving their experience.
[1364] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1365] Step 1:
[1366] The user puts on the device ("Smart Payment Assist Glasses") and the system starts automatically. The input is the user's wearing information. The device activates the camera and gaze measurement means and begins to capture the user's gaze information and facial expressions in real time. The output obtained is gaze data and facial expression data captured in real time.
[1367] Step 2:
[1368] The device uses the front camera to capture an image of the product page the user is viewing. The input is the product page the user is viewing. The image data captured by the camera is sent to the server. The output is the image data sent to the server.
[1369] Step 3:
[1370] The server analyzes the received image data using OCR (Optical Character Recognition) technology to extract text data related to the product. The input is image data captured by a camera. Through OCR analysis, text data is extracted to obtain the specific product information the user is viewing. The output is the extracted text data.
[1371] Step 4:
[1372] The device uses gaze measurement means to collect the user's gaze information in real time. The input is the user's eye movement and focus. The gaze data is continuously sent to a server and used to identify the specific location the user is looking at. The output is the identified gaze information.
[1373] Step 5:
[1374] The server uses an emotion engine to analyze the user's facial expressions and tone of voice to obtain current emotion data. The inputs are the user's facial expressions and tone of voice. The emotion data is analyzed along with the gaze data to help identify the user's current emotional state. The output is the identified emotion data.
[1375] Step 6:
[1376] The server comprehensively analyzes gaze data, text data, and emotion data to determine the products in which the user is interested and the degree of their willingness to purchase them. The inputs are gaze data, text data, and emotion data. The analysis results in the recommendation of appropriate products and services. The output is information about recommended products and services.
[1377] Step 7:
[1378] Based on the analyzed content, the server selects appropriate background music using a background music selection means. The inputs are the user's emotional state and the text content being viewed. The selected music is sent to the terminal and played to the user through a bone conduction speaker. The resulting output is the selected background music.
[1379] Step 8:
[1380] The server generates explanations of difficult words and contexts for the user and converts them into audio data. The input is the text content the user is viewing and the identified difficult parts. The explanation generation means generates a summary and sends it to the terminal as audio data. The output obtained is explanatory information as audio data.
[1381] Step 9:
[1382] The server proposes the optimal payment method based on the analyzed emotional data and purchasing intention information and notifies the user. The input is the user's emotional state and recommended product information. The payment assistance means selects the appropriate payment method and provides assistance to the user. The output obtained is the proposed payment method.
[1383] This allows users to receive the best music and explanations, product recommendations, and payment assistance while reading or shopping online.
[1384] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1385] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1386] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1387] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1388] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1389] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1390] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1391] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1392] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1393] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1394] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1395] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1396] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1397] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1398] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1399] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1400] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1401] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1402] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1403] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1404] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1405] The following is further disclosed regarding the above embodiment.
[1406] (Claim 1)
[1407] A means for detecting the reader's gaze direction using a camera installed in front of the device;
[1408] A gaze measurement method that collects reader gaze information in real time,
[1409] An artificial intelligence means for analyzing the content of what is being read based on collected gaze information and data from the front camera;
[1410] background music selection means for selecting appropriate background music based on the analyzed content;
[1411] bone conduction speaker means for playing selected background music;
[1412] The system includes an explanation generator for generating and providing explanations of difficult words and contexts as needed.
[1413] (Claim 2)
[1414] 2. The system according to claim 1, wherein the background music selection means selects background music based on the emotion or theme of the analyzed content.
[1415] (Claim 3)
[1416] 2. The system according to claim 1, wherein the explanation generating means has a function of summarizing the meaning of the text and converting it into audio to provide it.
[1417] "Example 1"
[1418] (Claim 1)
[1419] a means for detecting the direction of a user's line of sight using an imaging device installed in front of the user;
[1420] A gaze tracking means for collecting user gaze information in real time;
[1421] A machine learning method that analyzes the content of reading based on collected gaze information and data from a photographing device;
[1422] background sound selection means for selecting appropriate background sound based on the analyzed content;
[1423] bone conduction sound means for reproducing selected background sounds;
[1424] The system includes an explanation generator for generating and providing explanations of difficult words and contexts as needed.
[1425] (Claim 2)
[1426] 2. The system according to claim 1, wherein the background sound selection means selects the background sound based on the emotion or theme of the analyzed content.
[1427] (Claim 3)
[1428] 2. The system according to claim 1, wherein the explanation generating means has a function of summarizing the meaning of the text and converting it into audio to provide it.
[1429] "Application Example 1"
[1430] (Claim 1)
[1431] A means for detecting the reader's gaze direction using a camera installed in front of the device;
[1432] A gaze measurement method that collects reader gaze information in real time,
[1433] An artificial intelligence means for analyzing the content of what is being read based on collected gaze information and data from the front camera;
[1434] background music selection means for selecting appropriate background music based on the analyzed content;
[1435] a sound wave conducting acoustic device for playing selected background music;
[1436] an explanation generation means for generating and providing explanations of difficult words and contexts as needed;
[1437] a data communication means for transmitting text and eye gaze information during reading to a data server, for the server to analyze the transmitted data, and for providing the analysis results to the user;
[1438] an audio control means for transmitting background music selected based on the analysis result to the smart device and for the smart device to play the background music;
[1439] A system including:
[1440] (Claim 2)
[1441] 2. The system according to claim 1, wherein the background music selection means selects background music based on the emotion or theme of the analyzed content.
[1442] (Claim 3)
[1443] 2. The system according to claim 1, wherein the explanation generating means has a function of summarizing the meaning of the text and converting it into audio to provide it.
[1444] (Claim 4)
[1445] The system according to claim 1, further comprising a data communication means for transmitting gaze information to a cloud server while a user is reading and analyzing the data.
[1446] (Claim 5)
[1447] The system according to claim 1, further comprising an acoustic control means for transmitting background music selected based on the analysis results to a smartphone and playing it on an acoustic wave conducting acoustic device.
[1448] "Example 2: Combining Emotion Engines"
[1449] (Claim 1)
[1450] A means for detecting the reader's line of sight using an image sensor installed in front of the device;
[1451] A gaze measurement device that collects reader gaze information in real time,
[1452] an artificial intelligence means for analyzing the content of the reading based on the collected gaze information and data from the forward image sensor;
[1453] background music selection means for selecting appropriate background music based on the analyzed content and reader's emotional data;
[1454] a bone conduction audio device for playing selected background music;
[1455] an explanation generation means for generating and providing explanations of difficult terms and contexts as needed;
[1456] The system includes a voice conversion means for converting the analyzed data into voice data and providing it to a reader.
[1457] (Claim 2)
[1458] 2. The system according to claim 1, wherein the background music selection means selects background music based on the emotion or theme of the analyzed content.
[1459] (Claim 3)
[1460] The system of claim 1, wherein the explanation generation means has the function of summarizing the meaning of the text using a generative AI model and converting it into audio to provide.
[1461] "Application example 2 when combining emotion engines"
[1462] (Claim 1)
[1463] A means for detecting a user's line of sight using a camera installed in front of the device;
[1464] A gaze measurement means for collecting user gaze information in real time;
[1465] An artificial intelligence means for analyzing the content of reading and online activities based on collected gaze information and data from the front camera;
[1466] background music selection means for selecting appropriate background music based on the analyzed content;
[1467] bone conduction speaker means for playing selected background music;
[1468] an explanation generation means for generating and providing explanations of difficult words and contexts as needed;
[1469] A product recommendation method that captures the product page that a user is viewing while online shopping and recommends appropriate products and services based on that information and emotional data;
[1470] The system includes a payment assistance means that suggests the optimal payment method based on analyzed emotional data and supports purchasing behavior.
[1471] (Claim 2)
[1472] 2. The system according to claim 1, wherein the background music selection means selects background music based on the emotion or theme of the analyzed content.
[1473] (Claim 3)
[1474] 2. The system according to claim 1, wherein the explanation generating means has a function of summarizing the meaning of the text and converting it into audio to provide it. [Explanation of symbols]
[1475] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for detecting the reader's gaze direction using a camera installed in front of the device; A gaze measurement method that collects reader gaze information in real time, An artificial intelligence means for analyzing the content of what is being read based on collected gaze information and data from the front camera; background music selection means for selecting appropriate background music based on the analyzed content; bone conduction speaker means for playing selected background music; The system includes an explanation generator for generating and providing explanations of difficult words and contexts as needed.
2. 2. The system according to claim 1, wherein the background music selection means selects background music based on the emotion and theme of the analyzed content.
3. 2. The system according to claim 1, wherein the explanation generating means has a function of summarizing the meaning of the text and converting it into audio to provide it.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A