system

The system addresses dyslexia by converting text into visually understandable images or videos, tailored to individual user profiles and emotional states, facilitating easier comprehension.

JP2026037434APending Publication Date: 2026-03-06SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Users with dyslexia face difficulty in understanding text information due to reading and writing challenges, necessitating a system to convert text into visually easy-to-understand content.

Method used

A system that acquires text information, converts it into images or videos, customizes the content based on user profiles, and transmits it to visual devices like AR or VR devices for easier comprehension.

Benefits of technology

Enables users with dyslexia to visually understand textual information in daily life and studies by providing customized and emotionally optimized visual content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026037434000001_ABST
    Figure 2026037434000001_ABST
Patent Text Reader

Abstract

To provide a system that solves the problem of users with dyslexia (a condition that makes it difficult to read and write characters) having difficulty easily understanding text information in daily life, and reduces the burden on users by converting text information into content that is visually easy to understand and displaying it in a customized manner. [Solution] A system including a means for acquiring text information, a means for converting the acquired text information into an image or video, a means for customizing the image or video based on a user's profile information, and a means for transmitting the customized image or video to the user's visual device.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] This invention aims to solve the problem that users with dyslexia (a condition that makes it difficult for them to read and write) have difficulty easily understanding text information in their daily lives. Specifically, the purpose is to reduce the burden on users by converting text information from signs, road signs, textbooks, etc. into visually easy-to-understand content and customizing and displaying it. [Means for solving the problem]

[0005] The present invention solves the above problems by the following means.

[0006] A system is provided that includes a means for acquiring text information, a means for converting the acquired text information into an image or video, a means for customizing the image or video based on a user's profile information, and a means for transmitting the customized image or video to the user's visual device, thereby assisting a user with dyslexia in visually understanding text information in daily life and learning.

[0007] "Text information" refers to information or sentences written in characters, and is expressed as a specific string of characters.

[0008] "Means of acquisition" refers to the functionality or equipment for collecting information or data using specific devices or technologies.

[0009] "Means of converting into images or videos" refers to the technology or methods for converting text information into a visually easy-to-understand format (still images or video).

[0010] "User profile information" means individual information associated with a particular user (e.g., type of visual impairment, preferences, etc.) that serves as a basis for the system to provide appropriate customization.

[0011] "Customization means" refers to techniques and capabilities for tailoring or modifying generated visual content based on specific user requirements or profiles.

[0012] A "visual device" is an electronic device that allows a user to visually receive and display content, including, for example, an AR (augmented reality) device or a VR (virtual reality) device. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] The system of the present invention is designed to help users with dyslexia visually understand textual information in their daily lives and studies. This system involves a series of processes in which the user's terminal acquires textual information, which is then converted into visual content by a server, which customizes the content, and then transmits it to a visual device.

[0035] System configuration and program handling

[0036] 1. Obtaining text information

[0037] Users use devices such as smartphones or AR glasses to capture the text information they want to read. Specifically, they can take a picture of a sign or a textbook page with a camera. The device then extracts the text information from the captured image using OCR (optical character recognition) technology. The extracted text information is then sent to a server via a communications network.

[0038] 2. Analyzing textual information and converting it into visual content

[0039] The server receives the text information sent from the device and performs syntactic analysis. Syntactic analysis allows the server to understand the meaning and structure of the text information. The server then uses a natural language processing (NLP) model to convert the text information into a visually understandable format. This converted text information is then further converted into images or videos using generative AI (e.g., GAN: Generative Artificial Intelligence Network).

[0040] 3. Individual customization

[0041] The server customizes the generated visual content based on the user's profile information (type of visual impairment, customization preferences, etc.) For example, it may emphasize certain colors for color-blind users, or adjust font size and placement to make them easier to understand. The customized visual content is temporarily stored in a database on the server.

[0042] 4. Transfer and display on a visual device

[0043] The server sends the customized visual content to the user's device or visual device (e.g., AR glasses or VR headset). The device displays the received visual content, allowing the user to visually confirm the text information in real time. This display makes it easy for users with disabilities, such as dyslexia, to understand the information.

[0044] Specific examples

[0045] When reading signs in the city

[0046] A user points their smartphone camera at a sign and captures it. The device extracts text from the sign image and sends it to the server. The server then analyzes the received text information and uses generative AI to convert the text into visual content. The server then customizes the visual content based on the user's profile information and sends it to the user's AR device. Finally, the user can visually understand the information on the sign through the AR device.

[0047] When reading the contents of a textbook

[0048] A student user takes a photo of a textbook page with their smartphone. The device extracts text from the textbook image and sends it to the server. The server then analyzes the received text information and converts it into a visually understandable format. The server then customizes it based on the user's profile information and sends the customized visual content to the user's device. The user can then view the textbook content in an easy-to-understand format through their visual device.

[0049] In this way, the system of the present invention provides a new method for making textual information easier to understand visually for users with dyslexia.

[0050] The processing flow will be explained below.

[0051] Step 1:

[0052] A user uses a device such as a smartphone or AR glasses to capture text information to be read. For example, the user takes a picture of a sign or a page in a textbook with a camera. The device then acquires the captured image data.

[0053] Step 2:

[0054] The device extracts text information from the acquired image data using OCR (Optical Character Recognition) technology, which converts the characters in the image into digital text. The device then sends the extracted text information to the server.

[0055] Step 3:

[0056] The server receives the text information sent from the terminal, analyzes the received text information, and performs syntax analysis. This analysis allows the meaning and structure of the text to be understood.

[0057] Step 4:

[0058] The server uses a natural language processing (NLP) model to convert the parsed text information into a visually understandable format, and then invokes a generative AI (e.g., GAN: Generative Artificial Intelligence Network) to convert the text into visual content such as images or videos, thereby generating the text information in a visual format.

[0059] Step 5:

[0060] The server references the user's profile information, which includes the user's type of visual impairment and customization preferences, and customizes the generated visual content based on this information. For example, it may enhance certain colors for users with color blindness, or change font size or layout.

[0061] Step 6:

[0062] The server then sends the customized visual content to the user's device or visual device (such as AR glasses or VR headset), in real time, allowing the user to view the information instantly.

[0063] Step 7:

[0064] The device displays the received visual content. By looking at the visual content generated through the device or visual device, the user can visually understand the text information they have read. This makes it easier for people with disabilities, such as dyslexia, to understand the information.

[0065] Example 1

[0066] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0067] Users with dyslexia have difficulty visually understanding textual information in their daily lives and in their studies. This problem is particularly pronounced when reading books or recognizing street signs. Conventional methods lack an efficient system for converting textual information into a visually understandable format. Given this background, there is a need to provide a method that enables users with dyslexia to more easily visually understand textual information.

[0068] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0069] In this invention, the server includes means for parsing text information and converting it into visual content, means for converting it into a visually understandable format using a generative AI model, and means for customizing the visual content based on user profile information, thereby enabling users with dyslexia to convert and display text information in a visually understandable format in real time.

[0070] "Text information" refers to information expressed as characters or sentences.

[0071] A "communications network" is an infrastructure for transmitting and receiving data.

[0072] A "server" is a computer system that processes and stores data on a network.

[0073] "Syntax analysis" is the process of understanding the grammar and structure of given text data.

[0074] "Visual content" refers to media that is displayed visually, such as images and videos.

[0075] A "generative AI model" is a learning algorithm for generating data using artificial intelligence techniques.

[0076] "User profile information" means information about an individual user, such as settings, preferences, and characteristics.

[0077] "Customization" refers to changing settings and content to suit individual needs and preferences.

[0078] A "visual device" is a device for displaying visual information, including AR devices and VR devices.

[0079] "Optical character recognition (OCR) technology" is a technology that extracts character information from image data.

[0080] This invention is a system for supporting users with dyslexia to visually understand textual information in their daily lives and studies. This system involves a series of processes: a terminal acquires textual information, a server converts it into visual content, customizes it, and sends it to a visual device.

[0081] First, a user captures the text information they want to read using a device such as a smartphone or AR glasses. Specifically, the user uses the smartphone camera to take a picture of a textbook page or a sign in the city. The device then uses OCR (optical character recognition) technology to extract the text information from the captured image. For example, Google® Cloud Vision API is used for this extraction.

[0082] The terminal then transmits the extracted text information to a server over a communications network, typically using an HTTP POST request, with an endpoint configured using, for example, AWS® API Gateway.

[0083] The server first parses the received text information. For example, the server uses the Python library spaCy to analyze the structure and meaning of the text. The server then uses an NLP (natural language processing) model to convert the text information into a form that is easy to understand visually. Specifically, OpenAI (registered trademark) generative AI models (e.g., GPT-3 (registered trademark)) or generative artificial network (GAN) models are used.

[0084] Additionally, the server customizes the generated visual content based on the user's profile information, for example, increasing font size or enhancing colors for a user with dyslexia, based on the user's preference database.

[0085] Finally, the server sends the customized visual content to the user's device or visual device (e.g., AR glasses or VR headset), again using an HTTP POST request. The device displays the received visual content, allowing the user to visually confirm text information in real time. This helps users with dyslexia to more easily understand text information in their daily lives.

[0086] Specific examples

[0087] When reading signs in the city

[0088] The user points their smartphone camera at a sign and captures it. The device uses the Google Cloud Vision API to extract text from the sign image and sends it to the server via AWS API Gateway. The server then analyzes the received text using spaCy and converts it into a visually understandable format using a generative AI model. The text is then further customized based on the user's profile information, and the final visual content is sent to the user's AR device. Through the AR device, the user can visually perceive the sign's information, for example, with larger, more emphasized text.

[0089] When reading the contents of a textbook

[0090] A student user takes a photo of a textbook page with their smartphone camera. The device uses the Google Cloud Vision API to extract text information from the textbook image and sends it to a server using AWS API Gateway. The server analyzes the received text information and converts it into visual content that is easy to understand using OpenAI's GPT-3 or GAN. This content is customized based on the user's profile information and then sent to the vision device. Finally, the user can view the textbook content on the vision device with text alignment and font size adjusted.

[0091] Prompt Sentence Examples

[0092] An example of a prompt in a generative AI model is, "Please convert this text information into a visually easy-to-understand image. User profile information is as follows: type of visual impairment is color blindness, customization preference is to set text background color to blue."

[0093] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0094] Step 1:

[0095] Users capture text information using devices such as smartphones or AR glasses. For example, they take a picture of a page in a textbook or a sign in the city using the smartphone camera.

[0096] Input: Physical text information (textbook pages, signs, etc.)

[0097] Output: Captured image data

[0098] Step 2:

[0099] The device uses OCR technology on the captured image to extract text information, using the Google Cloud Vision API, and then converts the extracted text information into a data format.

[0100] Input: Photographed image data

[0101] Output: Extracted text information data

[0102] Step 3:

[0103] The device sends the extracted text information to a server via a communication network. The device securely transfers the data to the server using an HTTP POST request, for example, using AWS API Gateway.

[0104] Input: Extracted text information data

[0105] Output: Text information sent to the server

[0106] Step 4:

[0107] The server receives the text information sent from the terminal and performs syntax analysis. The server uses a Python library (e.g., spaCy) to analyze the structure of the text and understand its grammar and meaning.

[0108] Input: Text information sent to the server

[0109] Output: Structure data of parsed text information

[0110] Step 5:

[0111] The server uses natural language processing (NLP) models to convert text information into a visually understandable format. The server then uses generative AI models (e.g., GPT-3) or GANs to convert the analyzed text information into visual content (images and videos).

[0112] Input: Structure data of parsed text information

[0113] Output: Visual content that is easy to understand

[0114] Step 6:

[0115] The server customizes the generated visual content based on the user's profile information, such as increasing font size or highlighting certain colors for a user with dyslexia, based on the user's preference database.

[0116] Input: User profile information, visual content

[0117] Output: Customized visual content

[0118] Step 7:

[0119] The server sends the customized visual content to the user's device or viewing device, again using an HTTP POST request to transfer the data. The device then displays the received visual content.

[0120] Input: Customized visual content

[0121] Output: Visual content displayed on the device

[0122] Step 8:

[0123] Users can see the received visual content in real time, making it easier for users with disabilities such as dyslexia to understand the information.

[0124] Input: Visual content displayed on the device

[0125] Output: Visually understood text information

[0126] (Application example 1)

[0127] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0128] Users with dyslexia have difficulty visually understanding written information such as product labels and guide signs in physical stores. This makes it difficult for users to obtain appropriate information, limiting their ability to select products and use store facilities. There is a need for a system that can support such users and enable them to shop and use physical stores comfortably.

[0129] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0130] In this invention, the server includes means for capturing product labels and guide signs in a physical store using the smart glasses and converting the text information into a visually understandable format, means for customizing the images or videos based on the user's profile information, and means for displaying the generated visual content on the smart glasses so that the user can view the visual information in real time, thereby enabling users with dyslexia to easily understand products and guide signs in a physical store and enjoy shopping and facility use.

[0131] "Text information" refers to character data found in books, signs, product labels, and the like.

[0132] "Image or video" refers to still images or dynamic video data that visually represent text information.

[0133] "User profile information" refers to personalized information such as the type of visual impairment the user has and their customization preferences.

[0134] A "visual device" is a device that allows a user to receive visual information, specifically smart glasses, augmented reality (AR) devices, and virtual reality (VR) devices.

[0135] A "brick and mortar store" is a physical location for selling goods or providing services.

[0136] "Capture" refers to the act of using a device such as a camera to obtain images of product labels or guide signs in a physical store.

[0137] "Optical character recognition (OCR) technology" is a technology for automatically extracting text information from images.

[0138] "Smart glasses" are eyeglass-type devices equipped with camera and display functions, allowing users to check visual information in real time.

[0139] This invention is a system for supporting users with dyslexia to visually understand text information such as product labels and guide signs in physical stores. A specific embodiment of this system is described below.

[0140] First, a user wears smart glasses and captures product labels and guide signs while walking around a physical store. The smart glasses have a camera function and extract text information from the captured images using optical character recognition (OCR) technology (e.g., Tesseract OCR). The extracted text information is then sent from the smart glasses to a server via a communication network.

[0141] The server then parses the received text information and uses a natural language processing (NLP) model (e.g., BERT, GPT-3) to understand its meaning and structure. The text information is converted into a visually understandable format, which is then transformed into visual content in the form of images or videos using a generative AI model (e.g., GAN, DALL-E).

[0142] The server then customizes the generated visual content based on the user's profile information (e.g., type of visual impairment, customization preferences, etc.), for example, by highlighting certain colors or adjusting font size. The customized visual content is temporarily stored in a database within the server.

[0143] Finally, the server transmits the customized visual content to the smart glasses via a communication network, and the smart glasses display the received visual content for the user to view in real time.

[0144] As a concrete example, consider a user with dyslexia capturing a product label in the detergent section of a brick-and-mortar store. The user uses the camera in their smart glasses to capture the label: "Detergent - 750 yen special offer!"

[0145] 1. The camera takes a picture of the label and uses OCR technology (e.g., Tesseract OCR) to extract the text information: "Detergent - Special Price: 750 yen!"

[0146] 2. This text information is sent from the smart glasses to the server.

[0147] 3. The server uses an NLP model (e.g., GPT-3) to convert the text information into a visually understandable format, and a generative AI model (e.g., GAN) to generate a customized image (e.g., highlighting the "On Sale!" part in red).

[0148] 4. The server sends the generated visual content to the smart glasses, and the user can view the received visual content in real time.

[0149] A specific example of a prompt is "Please convert the following text into a visually understandable format. Text: 'Detergent - Special Offer for 750 yen!'" Based on this prompt, the server converts the text information using an NLP model and then generates visual content that is easy to understand using a generative AI model.

[0150] This allows users with dyslexia to easily understand products and information in physical stores, allowing them to shop and use facilities comfortably.

[0151] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0152] Step 1:

[0153] The user wears smart glasses and uses the camera to capture product labels and information signs in a physical store. The input is image data captured by the smart glasses' camera. The output is saved as image data on the device. Specifically, the user points the camera at the information they want to see and presses a button to capture the image.

[0154] Step 2:

[0155] The device extracts text information from the captured image using OCR (Optical Character Recognition) technology (e.g., Tesseract OCR). The input is the captured image data. The output is the extracted text information. Specifically, the OCR software analyzes the image, recognizes the text, and extracts it as character data.

[0156] Step 3:

[0157] The extracted text information is sent from the device to the server via a communication network. The input is the text information extracted by the OCR. The output is the text information received by the server. Specifically, the device uploads the text information to the server via an Internet connection.

[0158] Step 4:

[0159] The server uses an NLP model (e.g., BERT, GPT-3) to parse the received text information. The input is the text information sent from the device. The output is the text information converted to make it easier to understand visually. Specifically, the NLP model analyzes the meaning and structure of the text and performs natural language processing.

[0160] Step 5:

[0161] The server uses a generative AI model (e.g., GAN, DALL-E) to generate images and videos based on the information obtained using the NLP model and converts it into a visually easy-to-understand format. The input is text information analyzed by the NLP model. The output is images or videos in a visually easy-to-understand format. Specifically, the generative AI model creates appropriate visual content.

[0162] Step 6:

[0163] The server customizes the generated visual content based on the user's profile information. The input is images or videos converted to be visually easy to understand, along with the user's profile information. The output is customized visual content. Specific actions include customization to suit the user's needs, such as enhancing colors or adjusting font sizes.

[0164] Step 7:

[0165] The server transmits the customized visual content to the user's smart glasses via a communication network. The input is the customized visual content. The output is the visual content received by the smart glasses. Specific operations include data transfer from the server to the smart glasses.

[0166] Step 8:

[0167] The terminal displays the received visual content, and the user can view it in real time. The input is the customized visual content sent from the server. The output is the information that the user visually views. In concrete terms, the visual content is displayed on the display of the smart glasses, allowing the user to view the information.

[0168] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0169] The system of the present invention supports users with dyslexia in visually understanding textual information in their daily lives and studies. This embodiment also adds a function to recognize the user's emotions and optimize visual content based on those emotions. This system involves a series of processes: the user's device acquires textual information, the server converts it into customized visual content, and the server further optimizes it based on the user's emotions and transmits it to the visual device.

[0170] System configuration and program handling

[0171] 1. Obtaining text information

[0172] A user uses a device such as a smartphone or AR glasses to capture text information to be read. For example, they take a picture of a sign or a page in a textbook with a camera. The device then extracts the text information from the captured image using OCR (optical character recognition) technology. The extracted text information is then sent to a server via a communication network.

[0173] 2. Analyzing textual information and converting it into visual content

[0174] The server receives the text information sent from the device and performs syntactic analysis. Syntactic analysis understands the meaning and structure of the text. The server then uses a natural language processing (NLP) model to convert the text information into a visually understandable format. This converted text information is then further converted into images or videos using generative AI (e.g., GAN: Generative Artificial Intelligence Network).

[0175] 3. Individual customization

[0176] The server customizes the generated visual content based on the user's profile information (type of visual impairment, customization preferences, etc.) For example, it may emphasize certain colors for color-blind users, or adjust font size and placement to make them easier to understand. The customized visual content is temporarily stored in a database on the server.

[0177] 4. Optimization by Emotion Engine

[0178] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's facial expression data or voice data to identify the emotion the user is currently feeling (e.g., joy, sadness, surprise, anger, etc.). Based on this emotion data, the server further adjusts the customization of the visual content. For example, if the user is feeling stressed, the server may change the color tone to one that is visually more relaxing.

[0179] 5. Transfer and display on a visual device

[0180] The server sends the customized visual content to the user's device or visual device (e.g., AR glasses or VR headset). This transfer occurs in real time, allowing the user to instantly view the information. The device then displays the received visual content, allowing the user to visually interpret the text information.

[0181] Specific examples

[0182] When reading signs in the city

[0183] The user points their smartphone camera at a sign and captures it. The device extracts text from the sign image and sends it to the server. The server analyzes the received text information and uses generative AI to convert the text into visual content. The server then customizes the content based on the user's profile information and uses an emotion engine to recognize and optimize the user's emotional state. The customized visual content is then sent to the device, allowing the user to visually understand information about the city.

[0184] When reading the contents of a textbook

[0185] A student user takes a photo of a textbook page with their smartphone. The device extracts text from the textbook image and sends it to the server. The server analyzes the received text information and converts it into a visually understandable format. The server then customizes the content based on the user's profile information, and an emotion engine analyzes the user's concentration level and emotional state to optimize the visual content. The user can view the textbook content in an easy-to-understand format through the adjusted visual content.

[0186] In this way, the system of the present invention not only makes it easier for users with dyslexia to visually understand textual information, but also provides visual content that is optimally adjusted according to the user's emotional state.

[0187] The processing flow will be explained below.

[0188] Step 1:

[0189] A user uses a device such as a smartphone or AR glasses to capture text information to be read. For example, the user takes a picture of a sign or a page in a textbook with a camera. The device then acquires the captured image data.

[0190] Step 2:

[0191] The device extracts text information from the acquired image data using OCR (Optical Character Recognition) technology, which converts the characters in the image into digital text. The device then sends the extracted text information to the server.

[0192] Step 3:

[0193] The server receives the text information sent from the terminal, analyzes the received text information, and performs syntax analysis. This analysis allows the meaning and structure of the text to be understood.

[0194] Step 4:

[0195] The server uses a natural language processing (NLP) model to convert the parsed text information into a visually understandable format, and then invokes a generative AI (e.g., GAN: Generative Artificial Intelligence Network) to convert the text into visual content such as images or videos, thereby generating the text information in a visual format.

[0196] Step 5:

[0197] The server references the user's profile information, which includes the user's type of visual impairment and customization preferences, and customizes the generated visual content based on this information. For example, it may enhance certain colors for users with color blindness, or change font size or layout.

[0198] Step 6:

[0199] The server recognizes the user's emotions using an emotion engine, which analyzes the user's facial expression data or voice data to identify the emotion the user is currently feeling (e.g., joy, sadness, surprise, anger, etc.).

[0200] Step 7:

[0201] The server further adjusts the visual content based on the user's emotional state, for example, changing the color tones to be more visually relaxing if the user is feeling stressed.

[0202] Step 8:

[0203] The server then sends the customized visual content to the user's device or viewing device (e.g., AR glasses or VR headset). This transfer occurs in real time, allowing the user to view the information instantly.

[0204] Step 9:

[0205] The device displays the received visual content. By looking at the visual content generated through the device or visual device, the user can visually understand the text information they have read. This makes it easier for people with disabilities, such as dyslexia, to understand the information.

[0206] Example 2

[0207] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0208] Conventional visual content generation systems focus on converting text information into something visually easy to understand, but do not fully consider the visual characteristics and emotional state of each user. This makes it difficult to provide optimal visual information, especially for users with dyslexia. Furthermore, they lack flexible customization to allow users to obtain information without feeling stressed or burdened.

[0209] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring text information, means for converting the text information into images or videos, means for customizing the images or videos based on the user's profile information, means for analyzing the user's emotional state and optimizing the visual content, and means for transmitting the customized and optimized images or videos to the user's visual device. This makes it possible to provide optimized visual content that takes into account the visual characteristics and emotional state of each user.

[0210] "Text information" is data of characters and sentences that the user should read.

[0211] The "means for acquiring" is a mechanism for acquiring text information using the user's device.

[0212] The "means for converting into images or videos" is a mechanism for converting acquired text information into images or videos that are visually easy to understand.

[0213] "User profile information" is data about an individual user's visual characteristics, preferences, and disabilities.

[0214] A "means for customizing" is a mechanism for optimizing and tailoring generated visual content to a particular user based on the user's profile information.

[0215] The "emotional state of the user" refers to the mood or state of mind of the user that can be understood through detected facial expressions and voice data.

[0216] The "analyzing means" is a mechanism for analyzing facial expression data and voice data to understand the emotional state of the user.

[0217] The "optimizing means" is a mechanism that further adjusts the visual content based on the user's emotional state to improve user comfort.

[0218] A "visual device" is a device used by a user to view visual content, including, for example, an augmented reality device or a virtual reality device.

[0219] A "transmitting means" is a mechanism that transfers customized and optimized visual content to a user's visual device.

[0220] A "system" is a collection of devices and software that include each of these means and operate in conjunction with one another.

[0221] The system of the present invention is designed to help users with dyslexia visually understand written information in their daily lives and studies. This system involves a series of processes: acquiring the user's text information, analyzing and converting it on the server, and further customizing and optimizing it based on the user's profile information and emotional state.

[0222] Specific examples of hardware and software used

[0223] Hardware:

[0224] Smartphone

[0225] Augmented Reality (AR) Devices

[0226] Virtual Reality (VR) Devices

[0227] software:

[0228] OCR (optical character recognition) technology

[0229] Parsing Engine

[0230] Natural Language Processing (NLP) Models

[0231] Generative AI (e.g., GAN: Generative Adversarial Networks)

[0232] Sentiment Analysis Engine

[0233] System operation example

[0234] When reading signs in the city

[0235] The user points their smartphone camera at a sign and captures it. The device uses OCR technology to extract text information, which is then sent to the server. The server then parses the received text and uses generative AI to convert it into visually understandable images or videos. The server then customizes the visual content based on the user's profile information and optimizes it using an emotion engine to recognize the user's emotional state. The customized and optimized visual content is then sent to the device, allowing the user to instantly visually understand the information on signs around town.

[0236] When reading the contents of a textbook

[0237] A student user takes a photo of a textbook page with their smartphone. The device uses OCR technology to extract text from the textbook image and sends it to the server. The server then parses the received text information and uses generative AI to convert it into a visually understandable format. The server then customizes the visual content based on the user's profile information and optimizes it using an emotion engine to analyze the user's concentration level and emotional state. The user can visually understand the content of the textbook through the tailored visual content.

[0238] Prompt Sentence Examples

[0239] "Use GANs to run a program that translates text information into visual content for dyslexic users, and optimize it for the user's emotional state. For example, if the user is stressed, adjust the color tones to be more relaxing."

[0240] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0241] Step 1: Obtaining text information

[0242] A user uses a device such as a smartphone or AR glasses to capture text information using the camera. For example, the user takes a picture of a sign in the city or a page in a textbook.

[0243] Input: Image data taken by the user

[0244] The device inputs the image data into OCR (optical character recognition) software to extract text information.

[0245] Output: Extracted text information

[0246] The terminal compresses the extracted text information, performs an error check, encrypts it, and transmits it to the server.

[0247] Step 2: Parsing and transforming text information

[0248] The server receives the text information sent from the terminal and stores it in a database.

[0249] Input: Text information sent from the device

[0250] The server uses natural language processing (NLP) models to analyze the text and understand its meaning and context, including morphological analysis and dependency analysis.

[0251] Output: Parsed text information

[0252] The server uses generative AI (e.g., GAN) based on the analyzed text information to convert it into visually easy-to-understand image or video content.

[0253] Output: The generated visual content (images or videos)

[0254] Step 3: Individual customization

[0255] The server refers to the user's profile information (visual characteristics, preferences, etc.) stored in a database.

[0256] Input: Generated visual content, user profile information

[0257] The server customizes the visual content based on the user's profile information, for example by enhancing certain colors for color-blind users or adjusting font size and placement.

[0258] Output: Customized visual content

[0259] The server temporarily stores the customized visual content in a database.

[0260] Step 4: Optimizing with an Emotional Engine

[0261] The server inputs facial expression data and voice data acquired from the terminal into an emotion engine and analyzes the user's emotional state.

[0262] Input: facial expression data, voice data

[0263] The server uses an emotion engine to identify the user's emotion (joy, sadness, surprise, anger, etc.).

[0264] Output: Emotion analysis results

[0265] Based on the results of the emotion analysis, the server adjusts the color tone and layout of the visual content, optimizing it so that users can view information more comfortably.

[0266] Output: Optimized visual content

[0267] Step 5: Transfer and display to a visual device

[0268] The server transmits optimized visual content in real time to the user's device or visual device (such as AR glasses or VR headset).

[0269] Input: Optimized visual content

[0270] The terminal decodes the received visual content and displays it on its display.

[0271] Output: The displayed visual content

[0272] Users view customized and optimized visual content through their terminals and visual devices and visually understand textual information.

[0273] (Application example 2)

[0274] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0275] Traditionally, users with dyslexia have had difficulty visually understanding textual information. This can also hinder their ability to understand textual information during everyday activities such as online shopping and information searches, resulting in stress and misunderstandings. Furthermore, visual content provided without considering the user's emotional state can make comprehension even more difficult. The present invention aims to solve these problems and enable users with dyslexia to easily understand textual information and live their daily lives comfortably.

[0276] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring text information, means for analyzing the acquired text information and converting it into a form that is visually easy to understand, means for recognizing a user's emotion and optimizing visual content based on the emotion, means for customizing the visual content based on profile information, and means for transmitting the customized and optimized visual content to the user's terminal or visual device. This makes it easier for users with dyslexia to visually understand text information, and further enables the provision of optimal content according to their emotional state.

[0277] "Text information" is data of characters or sentences that are visually displayed.

[0278] "Means for acquiring text information" refers to a method of capturing visual characters or sentences as electronic data using a device such as a camera or scanner.

[0279] "Analysis" refers to processing the acquired text information to understand its meaning and structure.

[0280] "Means of converting into a visually easy-to-understand form" refers to methods of converting text information into a visually easy-to-understand form such as an image or video.

[0281] "Means for recognizing a user's emotions and optimizing visual content based on those emotions" refers to a method for analyzing a user's current emotional state from facial expressions and voice, and adjusting visual content to appropriately reflect those emotions.

[0282] "Profile information" refers to information such as the type of visual impairment and customization preferences of each individual user.

[0283] "Means for customizing visual content" refers to methods for adjusting color settings, font size, placement, etc. to suit each user based on profile information.

[0284] "Visual devices" are devices for visually displaying digital information, such as AR glasses and VR headsets.

[0285] A "server" is a centralized computer system for processing and managing data.

[0286] A "user's terminal" is a device that a user can operate at hand, such as a smartphone or tablet.

[0287] A "natural language processing model" is a learning model for analyzing and generating text data.

[0288] An "emotion engine" is software for analyzing a user's emotional state.

[0289] A "generative AI model" is an artificial intelligence model that uses neural networks to generate new data.

[0290] This invention provides a system that helps users with dyslexia visually understand text information in daily life and online shopping. In this embodiment, this system is realized through the following steps.

[0291] The server analyzes the text information sent by the user and converts it into a visually understandable format. First, the user captures the text information they want to obtain using a device such as a smartphone or tablet. From this captured image, the text information is extracted using optical character recognition (OCR) technology. The extracted text information is then sent to the server.

[0292] The server performs syntactic analysis on the received text to understand its meaning and structure, then uses a natural language processing (NLP) model to convert the text into a visually understandable format, which is then converted into an image or video using a generative AI model (e.g., a generative neural network).

[0293] The server then customizes the visual content based on the profile information, which may include, for example, the type of visual impairment and customization preferences, such as color enhancement or font size adjustment.

[0294] Additionally, the server uses an emotion engine to recognize the user's emotional state. It analyzes the user's facial expressions and voice data to identify their current emotion (happiness, sadness, stress, etc.). Based on the recognized emotion, the visual content is optimized. For example, if the user is feeling stressed, the visual color tone will be changed to a more relaxing one.

[0295] The customized and optimized visual content is then sent from the server to the user's terminal or visual device, such as an augmented reality (AR) glass or a virtual reality (VR) headset, which then displays the received visual content in real time, allowing the user to instantly visualize the information.

[0296] For example, when a user with dyslexia tries to understand product information on an online shopping website, the process is as follows: The user uses their smartphone camera to capture the product description. The device uses OCR technology to extract text from the image and sends it to the server. The server then analyzes the text and converts it into a visually understandable format. This converted information is customized based on the user's profile information and further optimized for the user's emotional state. The customized visual content is then sent to the user's device, where it can be transparently understood.

[0297] Examples of prompts:

[0298] "Please take a picture of this product description to make it easier to understand visually."

[0299] In this way, the system of the present invention can provide textual information in a visually easy-to-understand format when a user with dyslexia is shopping online, and can also provide visual content optimized according to the user's emotional state.

[0300] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0301] Step 1:

[0302] A user uses a smartphone or tablet to capture text information to read. Specifically, they open a camera app and take a picture of a product description or review. The input is the image taken with the smartphone camera, and the output is the captured image data.

[0303] Step 2:

[0304] The device extracts text information from the acquired image data using OCR technology. In this process, the image data is input, and an OCR engine (e.g., Tesseract) is used to recognize characters, resulting in text data as the output.

[0305] Step 3:

[0306] Text data is sent from a terminal to a server. The input is text data that is transferred to the server via a communication network. The output is the text data that arrives at the server.

[0307] Step 4:

[0308] The server analyzes the received text data and converts it into a form that is easy to understand visually. Specifically, it performs semantic analysis using a natural language processing model (e.g., BERT or GPT) and generates images and videos using a generative AI model (e.g., GAN). The input is text data and the output is visual content (images and videos).

[0309] Step 5:

[0310] The server customizes the visual content based on the user's profile information, for example by adjusting colors to accommodate color blindness or changing text size. The inputs to this process are the user's profile information and the visual content, and the output is the customized visual content.

[0311] Step 6:

[0312] The server uses an emotion engine to recognize the user's emotional state. Specifically, it analyzes facial expression images and voice data provided by the user to identify emotions. The input of this process is emotion data (facial expression images and voice data), and the output is the recognized emotional state.

[0313] Step 7:

[0314] The server further optimizes the visual content based on the recognized emotional state, for example, changing the color tones to a more relaxing color if the user is stressed. The inputs to this process are the emotional state and the customized visual content, and the output is the optimized visual content.

[0315] Step 8:

[0316] The optimized visual content is sent from the server to the user's device or visual device. The input is the optimized visual content, which is sent to the user's device or AR glasses via a communication network. The output is the visual content displayed on the user's device or visual device.

[0317] Step 9:

[0318] The user checks the visual content displayed on the terminal or visual device and visually understands the required information. The input is the optimized content displayed on the visual device, and the output is the user's understanding. This step allows the user to visually understand the information comfortably.

[0319] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0320] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0321] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0322] [Second embodiment]

[0323] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0324] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0325] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0326] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0327] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0328] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0329] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0330] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0331] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0332] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0333] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0334] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0335] The system of the present invention is designed to help users with dyslexia visually understand textual information in their daily lives and studies. This system involves a series of processes in which the user's terminal acquires textual information, which is then converted into visual content by a server, which customizes the content, and then transmits it to a visual device.

[0336] System configuration and program handling

[0337] 1. Obtaining text information

[0338] Users use devices such as smartphones or AR glasses to capture the text information they want to read. Specifically, they can take a picture of a sign or a textbook page with a camera. The device then extracts the text information from the captured image using OCR (optical character recognition) technology. The extracted text information is then sent to a server via a communications network.

[0339] 2. Analyzing textual information and converting it into visual content

[0340] The server receives the text information sent from the device and performs syntactic analysis. Syntactic analysis allows the server to understand the meaning and structure of the text information. The server then uses a natural language processing (NLP) model to convert the text information into a visually understandable format. This converted text information is then further converted into images or videos using generative AI (e.g., GAN: Generative Artificial Intelligence Network).

[0341] 3. Individual customization

[0342] The server customizes the generated visual content based on the user's profile information (type of visual impairment, customization preferences, etc.) For example, it may emphasize certain colors for color-blind users, or adjust font size and placement to make them easier to understand. The customized visual content is temporarily stored in a database on the server.

[0343] 4. Transfer and display on a visual device

[0344] The server sends the customized visual content to the user's device or visual device (e.g., AR glasses or VR headset). The device displays the received visual content, allowing the user to visually confirm the text information in real time. This display makes it easy for users with disabilities, such as dyslexia, to understand the information.

[0345] Specific examples

[0346] When reading signs in the city

[0347] A user points their smartphone camera at a sign and captures it. The device extracts text from the sign image and sends it to the server. The server then analyzes the received text information and uses generative AI to convert the text into visual content. The server then customizes the visual content based on the user's profile information and sends it to the user's AR device. Finally, the user can visually understand the information on the sign through the AR device.

[0348] When reading the contents of a textbook

[0349] A student user takes a photo of a textbook page with their smartphone. The device extracts text from the textbook image and sends it to the server. The server then analyzes the received text information and converts it into a visually understandable format. The server then customizes it based on the user's profile information and sends the customized visual content to the user's device. The user can then view the textbook content in an easy-to-understand format through their visual device.

[0350] In this way, the system of the present invention provides a new method for making textual information easier to understand visually for users with dyslexia.

[0351] The processing flow will be explained below.

[0352] Step 1:

[0353] A user uses a device such as a smartphone or AR glasses to capture text information to be read. For example, the user takes a picture of a sign or a page in a textbook with a camera. The device then acquires the captured image data.

[0354] Step 2:

[0355] The device extracts text information from the acquired image data using OCR (Optical Character Recognition) technology, which converts the characters in the image into digital text. The device then sends the extracted text information to the server.

[0356] Step 3:

[0357] The server receives the text information sent from the terminal, analyzes the received text information, and performs syntax analysis. This analysis allows the meaning and structure of the text to be understood.

[0358] Step 4:

[0359] The server uses a natural language processing (NLP) model to convert the parsed text information into a visually understandable format, and then invokes a generative AI (e.g., GAN: Generative Artificial Intelligence Network) to convert the text into visual content such as images or videos, thereby generating the text information in a visual format.

[0360] Step 5:

[0361] The server references the user's profile information, which includes the user's type of visual impairment and customization preferences, and customizes the generated visual content based on this information. For example, it may enhance certain colors for users with color blindness, or change font size or layout.

[0362] Step 6:

[0363] The server then sends the customized visual content to the user's device or visual device (such as AR glasses or VR headset), in real time, allowing the user to view the information instantly.

[0364] Step 7:

[0365] The device displays the received visual content. By looking at the visual content generated through the device or visual device, the user can visually understand the text information they have read. This makes it easier for people with disabilities, such as dyslexia, to understand the information.

[0366] Example 1

[0367] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0368] Users with dyslexia have difficulty visually understanding textual information in their daily lives and in their studies. This problem is particularly pronounced when reading books or recognizing street signs. Conventional methods lack an efficient system for converting textual information into a visually understandable format. Given this background, there is a need to provide a method that enables users with dyslexia to more easily visually understand textual information.

[0369] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0370] In this invention, the server includes means for parsing text information and converting it into visual content, means for converting it into a visually understandable format using a generative AI model, and means for customizing the visual content based on user profile information, thereby enabling users with dyslexia to convert and display text information in a visually understandable format in real time.

[0371] "Text information" refers to information expressed as characters or sentences.

[0372] A "communications network" is an infrastructure for transmitting and receiving data.

[0373] A "server" is a computer system that processes and stores data on a network.

[0374] "Syntax analysis" is the process of understanding the grammar and structure of given text data.

[0375] "Visual content" refers to media that is displayed visually, such as images and videos.

[0376] A "generative AI model" is a learning algorithm for generating data using artificial intelligence techniques.

[0377] "User profile information" means information about an individual user, such as settings, preferences, and characteristics.

[0378] "Customization" refers to changing settings and content to suit individual needs and preferences.

[0379] A "visual device" is a device for displaying visual information, including AR devices and VR devices.

[0380] "Optical character recognition (OCR) technology" is a technology that extracts character information from image data.

[0381] This invention is a system for supporting users with dyslexia to visually understand textual information in their daily lives and studies. This system involves a series of processes: a terminal acquires textual information, a server converts it into visual content, customizes it, and sends it to a visual device.

[0382] First, the user captures the text information they want to read using a device such as a smartphone or AR glasses. Specifically, the user uses the smartphone camera to take a picture of a textbook page or a sign in the city. The device then uses OCR (optical character recognition) technology to extract the text information from the captured image. This extraction can be done using, for example, the Google Cloud Vision API.

[0383] The device then sends the extracted text information to a server over a communications network, typically using an HTTP POST request, with an endpoint set up using, for example, AWS API Gateway.

[0384] The server first parses the received text information. For example, the server uses the Python library spaCy to analyze the structure and meaning of the text. The server then uses an NLP (natural language processing) model to convert the text information into a visually understandable format. Specifically, OpenAI's generative AI model (e.g., GPT-3) or a generative artificial network (GAN) model is used.

[0385] Additionally, the server customizes the generated visual content based on the user's profile information, for example, increasing font size or enhancing colors for a user with dyslexia, based on the user's preference database.

[0386] Finally, the server sends the customized visual content to the user's device or visual device (e.g., AR glasses or VR headset), again using an HTTP POST request. The device displays the received visual content, allowing the user to visually confirm text information in real time. This helps users with dyslexia to more easily understand text information in their daily lives.

[0387] Specific examples

[0388] When reading signs in the city

[0389] The user points their smartphone camera at a sign and captures it. The device uses the Google Cloud Vision API to extract text from the sign image and sends it to the server via AWS API Gateway. The server then analyzes the received text using spaCy and converts it into a visually understandable format using a generative AI model. The text is then further customized based on the user's profile information, and the final visual content is sent to the user's AR device. Through the AR device, the user can visually perceive the sign's information, for example, with larger, more emphasized text.

[0390] When reading the contents of a textbook

[0391] A student user takes a photo of a textbook page with their smartphone camera. The device uses the Google Cloud Vision API to extract text information from the textbook image and sends it to a server using AWS API Gateway. The server analyzes the received text information and converts it into visual content that is easy to understand using OpenAI's GPT-3 or GAN. This content is customized based on the user's profile information and then sent to the vision device. Finally, the user can view the textbook content on the vision device with text alignment and font size adjusted.

[0392] Prompt Sentence Examples

[0393] An example of a prompt in a generative AI model is, "Please convert this text information into a visually easy-to-understand image. User profile information is as follows: type of visual impairment is color blindness, customization preference is to set text background color to blue."

[0394] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0395] Step 1:

[0396] Users capture text information using devices such as smartphones or AR glasses. For example, they take a picture of a page in a textbook or a sign in the city using the smartphone camera.

[0397] Input: Physical text information (textbook pages, signs, etc.)

[0398] Output: Captured image data

[0399] Step 2:

[0400] The device uses OCR technology on the captured image to extract text information, using the Google Cloud Vision API, and then converts the extracted text information into a data format.

[0401] Input: Photographed image data

[0402] Output: Extracted text information data

[0403] Step 3:

[0404] The device sends the extracted text information to a server via a communication network. The device securely transfers the data to the server using an HTTP POST request, for example, using AWS API Gateway.

[0405] Input: Extracted text information data

[0406] Output: Text information sent to the server

[0407] Step 4:

[0408] The server receives the text information sent from the terminal and performs syntax analysis. The server uses a Python library (e.g., spaCy) to analyze the structure of the text and understand its grammar and meaning.

[0409] Input: Text information sent to the server

[0410] Output: Structure data of parsed text information

[0411] Step 5:

[0412] The server uses natural language processing (NLP) models to convert text information into a visually understandable format. The server then uses generative AI models (e.g., GPT-3) or GANs to convert the analyzed text information into visual content (images and videos).

[0413] Input: Structure data of parsed text information

[0414] Output: Visual content that is easy to understand

[0415] Step 6:

[0416] The server customizes the generated visual content based on the user's profile information, such as increasing font size or highlighting certain colors for a user with dyslexia, based on the user's preference database.

[0417] Input: User profile information, visual content

[0418] Output: Customized visual content

[0419] Step 7:

[0420] The server sends the customized visual content to the user's device or viewing device, again using an HTTP POST request to transfer the data. The device then displays the received visual content.

[0421] Input: Customized visual content

[0422] Output: Visual content displayed on the device

[0423] Step 8:

[0424] Users can see the received visual content in real time, making it easier for users with disabilities such as dyslexia to understand the information.

[0425] Input: Visual content displayed on the device

[0426] Output: Visually understood text information

[0427] (Application example 1)

[0428] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0429] Users with dyslexia have difficulty visually understanding written information such as product labels and guide signs in physical stores. This makes it difficult for users to obtain appropriate information, limiting their ability to select products and use store facilities. There is a need for a system that can support such users and enable them to shop and use physical stores comfortably.

[0430] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0431] In this invention, the server includes means for capturing product labels and guide signs in a physical store using the smart glasses and converting the text information into a visually understandable format, means for customizing the images or videos based on the user's profile information, and means for displaying the generated visual content on the smart glasses so that the user can view the visual information in real time, thereby enabling users with dyslexia to easily understand products and guide signs in a physical store and enjoy shopping and facility use.

[0432] "Text information" refers to character data found in books, signs, product labels, and the like.

[0433] "Image or video" refers to still images or dynamic video data that visually represent text information.

[0434] "User profile information" refers to personalized information such as the type of visual impairment the user has and their customization preferences.

[0435] A "visual device" is a device that allows a user to receive visual information, specifically smart glasses, augmented reality (AR) devices, and virtual reality (VR) devices.

[0436] A "brick and mortar store" is a physical location for selling goods or providing services.

[0437] "Capture" refers to the act of using a device such as a camera to obtain images of product labels or guide signs in a physical store.

[0438] "Optical character recognition (OCR) technology" is a technology for automatically extracting text information from images.

[0439] "Smart glasses" are eyeglass-type devices equipped with camera and display functions, allowing users to check visual information in real time.

[0440] This invention is a system for supporting users with dyslexia to visually understand text information such as product labels and guide signs in physical stores. A specific embodiment of this system is described below.

[0441] First, a user wears smart glasses and captures product labels and guide signs while walking around a physical store. The smart glasses have a camera function and extract text information from the captured images using optical character recognition (OCR) technology (e.g., Tesseract OCR). The extracted text information is then sent from the smart glasses to a server via a communication network.

[0442] The server then parses the received text information and uses a natural language processing (NLP) model (e.g., BERT, GPT-3) to understand its meaning and structure. The text information is converted into a visually understandable format, which is then transformed into visual content in the form of images or videos using a generative AI model (e.g., GAN, DALL-E).

[0443] The server then customizes the generated visual content based on the user's profile information (e.g., type of visual impairment, customization preferences, etc.), for example, by highlighting certain colors or adjusting font size. The customized visual content is temporarily stored in a database within the server.

[0444] Finally, the server transmits the customized visual content to the smart glasses via a communication network, and the smart glasses display the received visual content for the user to view in real time.

[0445] As a concrete example, consider a user with dyslexia capturing a product label in the detergent section of a brick-and-mortar store. The user uses the camera in their smart glasses to capture the label: "Detergent - 750 yen special offer!"

[0446] 1. The camera takes a picture of the label and uses OCR technology (e.g., Tesseract OCR) to extract the text information: "Detergent - Special Price: 750 yen!"

[0447] 2. This text information is sent from the smart glasses to the server.

[0448] 3. The server uses an NLP model (e.g., GPT-3) to convert the text information into a visually understandable format, and a generative AI model (e.g., GAN) to generate a customized image (e.g., highlighting the "On Sale!" part in red).

[0449] 4. The server sends the generated visual content to the smart glasses, and the user can view the received visual content in real time.

[0450] A specific example of a prompt is "Please convert the following text into a visually understandable format. Text: 'Detergent - Special Offer for 750 yen!'" Based on this prompt, the server converts the text information using an NLP model and then generates visual content that is easy to understand using a generative AI model.

[0451] This allows users with dyslexia to easily understand products and information in physical stores, allowing them to shop and use facilities comfortably.

[0452] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0453] Step 1:

[0454] The user wears smart glasses and uses the camera to capture product labels and information signs in a physical store. The input is image data captured by the smart glasses' camera. The output is saved as image data on the device. Specifically, the user points the camera at the information they want to see and presses a button to capture the image.

[0455] Step 2:

[0456] The device extracts text information from the captured image using OCR (Optical Character Recognition) technology (e.g., Tesseract OCR). The input is the captured image data. The output is the extracted text information. Specifically, the OCR software analyzes the image, recognizes the text, and extracts it as character data.

[0457] Step 3:

[0458] The extracted text information is sent from the device to the server via a communication network. The input is the text information extracted by the OCR. The output is the text information received by the server. Specifically, the device uploads the text information to the server via an Internet connection.

[0459] Step 4:

[0460] The server uses an NLP model (e.g., BERT, GPT-3) to parse the received text information. The input is the text information sent from the device. The output is the text information converted to make it easier to understand visually. Specifically, the NLP model analyzes the meaning and structure of the text and performs natural language processing.

[0461] Step 5:

[0462] The server uses a generative AI model (e.g., GAN, DALL-E) to generate images and videos based on the information obtained using the NLP model and converts it into a visually easy-to-understand format. The input is text information analyzed by the NLP model. The output is images or videos in a visually easy-to-understand format. Specifically, the generative AI model creates appropriate visual content.

[0463] Step 6:

[0464] The server customizes the generated visual content based on the user's profile information. The input is images or videos converted to be visually easy to understand, along with the user's profile information. The output is customized visual content. Specific actions include customization to suit the user's needs, such as enhancing colors or adjusting font sizes.

[0465] Step 7:

[0466] The server transmits the customized visual content to the user's smart glasses via a communication network. The input is the customized visual content. The output is the visual content received by the smart glasses. Specific operations include data transfer from the server to the smart glasses.

[0467] Step 8:

[0468] The terminal displays the received visual content, and the user can view it in real time. The input is the customized visual content sent from the server. The output is the information that the user visually views. In concrete terms, the visual content is displayed on the display of the smart glasses, allowing the user to view the information.

[0469] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0470] The system of the present invention supports users with dyslexia in visually understanding textual information in their daily lives and studies. This embodiment also adds a function to recognize the user's emotions and optimize visual content based on those emotions. This system involves a series of processes: the user's device acquires textual information, the server converts it into customized visual content, and the server further optimizes it based on the user's emotions and transmits it to the visual device.

[0471] System configuration and program handling

[0472] 1. Obtaining text information

[0473] A user uses a device such as a smartphone or AR glasses to capture text information to be read. For example, they take a picture of a sign or a page in a textbook with a camera. The device then extracts the text information from the captured image using OCR (optical character recognition) technology. The extracted text information is then sent to a server via a communication network.

[0474] 2. Analyzing textual information and converting it into visual content

[0475] The server receives the text information sent from the device and performs syntactic analysis. Syntactic analysis understands the meaning and structure of the text. The server then uses a natural language processing (NLP) model to convert the text information into a visually understandable format. This converted text information is then further converted into images or videos using generative AI (e.g., GAN: Generative Artificial Intelligence Network).

[0476] 3. Individual customization

[0477] The server customizes the generated visual content based on the user's profile information (type of visual impairment, customization preferences, etc.) For example, it may emphasize certain colors for color-blind users, or adjust font size and placement to make them easier to understand. The customized visual content is temporarily stored in a database on the server.

[0478] 4. Optimization by Emotion Engine

[0479] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's facial expression data or voice data to identify the emotion the user is currently feeling (e.g., joy, sadness, surprise, anger, etc.). Based on this emotion data, the server further adjusts the customization of the visual content. For example, if the user is feeling stressed, the server may change the color tone to one that is visually more relaxing.

[0480] 5. Transfer and display on a visual device

[0481] The server sends the customized visual content to the user's device or visual device (e.g., AR glasses or VR headset). This transfer occurs in real time, allowing the user to instantly view the information. The device then displays the received visual content, allowing the user to visually interpret the text information.

[0482] Specific examples

[0483] When reading signs in the city

[0484] The user points their smartphone camera at a sign and captures it. The device extracts text from the sign image and sends it to the server. The server analyzes the received text information and uses generative AI to convert the text into visual content. The server then customizes the content based on the user's profile information and uses an emotion engine to recognize and optimize the user's emotional state. The customized visual content is then sent to the device, allowing the user to visually understand information about the city.

[0485] When reading the contents of a textbook

[0486] A student user takes a photo of a textbook page with their smartphone. The device extracts text from the textbook image and sends it to the server. The server analyzes the received text information and converts it into a visually understandable format. The server then customizes the content based on the user's profile information, and an emotion engine analyzes the user's concentration level and emotional state to optimize the visual content. The user can view the textbook content in an easy-to-understand format through the adjusted visual content.

[0487] In this way, the system of the present invention not only makes it easier for users with dyslexia to visually understand textual information, but also provides visual content that is optimally adjusted according to the user's emotional state.

[0488] The processing flow will be explained below.

[0489] Step 1:

[0490] A user uses a device such as a smartphone or AR glasses to capture text information to be read. For example, the user takes a picture of a sign or a page in a textbook with a camera. The device then acquires the captured image data.

[0491] Step 2:

[0492] The device extracts text information from the acquired image data using OCR (Optical Character Recognition) technology, which converts the characters in the image into digital text. The device then sends the extracted text information to the server.

[0493] Step 3:

[0494] The server receives the text information sent from the terminal, analyzes the received text information, and performs syntax analysis. This analysis allows the meaning and structure of the text to be understood.

[0495] Step 4:

[0496] The server uses a natural language processing (NLP) model to convert the parsed text information into a visually understandable format, and then invokes a generative AI (e.g., GAN: Generative Artificial Intelligence Network) to convert the text into visual content such as images or videos, thereby generating the text information in a visual format.

[0497] Step 5:

[0498] The server references the user's profile information, which includes the user's type of visual impairment and customization preferences, and customizes the generated visual content based on this information. For example, it may enhance certain colors for users with color blindness, or change font size or layout.

[0499] Step 6:

[0500] The server recognizes the user's emotions using an emotion engine, which analyzes the user's facial expression data or voice data to identify the emotion the user is currently feeling (e.g., joy, sadness, surprise, anger, etc.).

[0501] Step 7:

[0502] The server further adjusts the visual content based on the user's emotional state, for example, changing the color tones to be more visually relaxing if the user is feeling stressed.

[0503] Step 8:

[0504] The server then sends the customized visual content to the user's device or viewing device (e.g., AR glasses or VR headset). This transfer occurs in real time, allowing the user to view the information instantly.

[0505] Step 9:

[0506] The device displays the received visual content. By looking at the visual content generated through the device or visual device, the user can visually understand the text information they have read. This makes it easier for people with disabilities, such as dyslexia, to understand the information.

[0507] Example 2

[0508] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0509] Conventional visual content generation systems focus on converting text information into something visually easy to understand, but do not fully consider the visual characteristics and emotional state of each user. This makes it difficult to provide optimal visual information, especially for users with dyslexia. Furthermore, they lack flexible customization to allow users to obtain information without feeling stressed or burdened.

[0510] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring text information, means for converting the text information into images or videos, means for customizing the images or videos based on the user's profile information, means for analyzing the user's emotional state and optimizing the visual content, and means for transmitting the customized and optimized images or videos to the user's visual device. This makes it possible to provide optimized visual content that takes into account the visual characteristics and emotional state of each user.

[0511] "Text information" is data of characters and sentences that the user should read.

[0512] The "means for acquiring" is a mechanism for acquiring text information using the user's device.

[0513] The "means for converting into images or videos" is a mechanism for converting acquired text information into images or videos that are visually easy to understand.

[0514] "User profile information" is data about an individual user's visual characteristics, preferences, and disabilities.

[0515] A "means for customizing" is a mechanism for optimizing and tailoring generated visual content to a particular user based on the user's profile information.

[0516] The "emotional state of the user" refers to the mood or state of mind of the user that can be understood through detected facial expressions and voice data.

[0517] The "analyzing means" is a mechanism for analyzing facial expression data and voice data to understand the emotional state of the user.

[0518] The "optimizing means" is a mechanism that further adjusts the visual content based on the user's emotional state to improve user comfort.

[0519] A "visual device" is a device used by a user to view visual content, including, for example, an augmented reality device or a virtual reality device.

[0520] A "transmitting means" is a mechanism that transfers customized and optimized visual content to a user's visual device.

[0521] A "system" is a collection of devices and software that include each of these means and operate in conjunction with one another.

[0522] The system of the present invention is designed to help users with dyslexia visually understand written information in their daily lives and studies. This system involves a series of processes: acquiring the user's text information, analyzing and converting it on the server, and further customizing and optimizing it based on the user's profile information and emotional state.

[0523] Specific examples of hardware and software used

[0524] Hardware:

[0525] Smartphone

[0526] Augmented Reality (AR) Devices

[0527] Virtual Reality (VR) Devices

[0528] software:

[0529] OCR (optical character recognition) technology

[0530] Parsing Engine

[0531] Natural Language Processing (NLP) Models

[0532] Generative AI (e.g., GAN: Generative Adversarial Networks)

[0533] Sentiment Analysis Engine

[0534] System operation example

[0535] When reading signs in the city

[0536] The user points their smartphone camera at a sign and captures it. The device uses OCR technology to extract text information, which is then sent to the server. The server then parses the received text and uses generative AI to convert it into visually understandable images or videos. The server then customizes the visual content based on the user's profile information and optimizes it using an emotion engine to recognize the user's emotional state. The customized and optimized visual content is then sent to the device, allowing the user to instantly visually understand the information on signs around town.

[0537] When reading the contents of a textbook

[0538] A student user takes a photo of a textbook page with their smartphone. The device uses OCR technology to extract text from the textbook image and sends it to the server. The server then parses the received text information and uses generative AI to convert it into a visually understandable format. The server then customizes the visual content based on the user's profile information and optimizes it using an emotion engine to analyze the user's concentration level and emotional state. The user can visually understand the content of the textbook through the tailored visual content.

[0539] Prompt Sentence Examples

[0540] "Use GANs to run a program that translates text information into visual content for dyslexic users, and optimize it for the user's emotional state. For example, if the user is stressed, adjust the color tones to be more relaxing."

[0541] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0542] Step 1: Obtaining text information

[0543] A user uses a device such as a smartphone or AR glasses to capture text information using the camera. For example, the user takes a picture of a sign in the city or a page in a textbook.

[0544] Input: Image data taken by the user

[0545] The device inputs the image data into OCR (optical character recognition) software to extract text information.

[0546] Output: Extracted text information

[0547] The terminal compresses the extracted text information, performs an error check, encrypts it, and transmits it to the server.

[0548] Step 2: Parsing and transforming text information

[0549] The server receives the text information sent from the terminal and stores it in a database.

[0550] Input: Text information sent from the device

[0551] The server uses natural language processing (NLP) models to analyze the text and understand its meaning and context, including morphological analysis and dependency analysis.

[0552] Output: Parsed text information

[0553] The server uses generative AI (e.g., GAN) based on the analyzed text information to convert it into visually easy-to-understand image or video content.

[0554] Output: The generated visual content (images or videos)

[0555] Step 3: Individual customization

[0556] The server refers to the user's profile information (visual characteristics, preferences, etc.) stored in a database.

[0557] Input: Generated visual content, user profile information

[0558] The server customizes the visual content based on the user's profile information, for example by enhancing certain colors for color-blind users or adjusting font size and placement.

[0559] Output: Customized visual content

[0560] The server temporarily stores the customized visual content in a database.

[0561] Step 4: Optimizing with an Emotional Engine

[0562] The server inputs facial expression data and voice data acquired from the terminal into an emotion engine and analyzes the user's emotional state.

[0563] Input: facial expression data, voice data

[0564] The server uses an emotion engine to identify the user's emotion (joy, sadness, surprise, anger, etc.).

[0565] Output: Emotion analysis results

[0566] Based on the results of the emotion analysis, the server adjusts the color tone and layout of the visual content, optimizing it so that users can view information more comfortably.

[0567] Output: Optimized visual content

[0568] Step 5: Transfer and display to a visual device

[0569] The server transmits optimized visual content in real time to the user's device or visual device (such as AR glasses or VR headset).

[0570] Input: Optimized visual content

[0571] The terminal decodes the received visual content and displays it on its display.

[0572] Output: The displayed visual content

[0573] Users view customized and optimized visual content through their terminals and visual devices and visually understand textual information.

[0574] (Application example 2)

[0575] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0576] Traditionally, users with dyslexia have had difficulty visually understanding textual information. This can also hinder their ability to understand textual information during everyday activities such as online shopping and information searches, resulting in stress and misunderstandings. Furthermore, visual content provided without considering the user's emotional state can make comprehension even more difficult. The present invention aims to solve these problems and enable users with dyslexia to easily understand textual information and live their daily lives comfortably.

[0577] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring text information, means for analyzing the acquired text information and converting it into a form that is visually easy to understand, means for recognizing a user's emotion and optimizing visual content based on the emotion, means for customizing the visual content based on profile information, and means for transmitting the customized and optimized visual content to the user's terminal or visual device. This makes it easier for users with dyslexia to visually understand text information, and further enables the provision of optimal content according to their emotional state.

[0578] "Text information" is data of characters or sentences that are visually displayed.

[0579] "Means for acquiring text information" refers to a method of capturing visual characters or sentences as electronic data using a device such as a camera or scanner.

[0580] "Analysis" refers to processing the acquired text information to understand its meaning and structure.

[0581] "Means of converting into a visually easy-to-understand form" refers to methods of converting text information into a visually easy-to-understand form such as an image or video.

[0582] "Means for recognizing a user's emotions and optimizing visual content based on those emotions" refers to a method for analyzing a user's current emotional state from facial expressions and voice, and adjusting visual content to appropriately reflect those emotions.

[0583] "Profile information" refers to information such as the type of visual impairment and customization preferences of each individual user.

[0584] "Means for customizing visual content" refers to methods for adjusting color settings, font size, placement, etc. to suit each user based on profile information.

[0585] "Visual devices" are devices for visually displaying digital information, such as AR glasses and VR headsets.

[0586] A "server" is a centralized computer system for processing and managing data.

[0587] A "user's terminal" is a device that a user can operate at hand, such as a smartphone or tablet.

[0588] A "natural language processing model" is a learning model for analyzing and generating text data.

[0589] An "emotion engine" is software for analyzing a user's emotional state.

[0590] A "generative AI model" is an artificial intelligence model that uses neural networks to generate new data.

[0591] This invention provides a system that helps users with dyslexia visually understand text information in daily life and online shopping. In this embodiment, this system is realized through the following steps.

[0592] The server analyzes the text information sent by the user and converts it into a visually understandable format. First, the user captures the text information they want to obtain using a device such as a smartphone or tablet. From this captured image, the text information is extracted using optical character recognition (OCR) technology. The extracted text information is then sent to the server.

[0593] The server performs syntactic analysis on the received text to understand its meaning and structure, then uses a natural language processing (NLP) model to convert the text into a visually understandable format, which is then converted into an image or video using a generative AI model (e.g., a generative neural network).

[0594] The server then customizes the visual content based on the profile information, which may include, for example, the type of visual impairment and customization preferences, such as color enhancement or font size adjustment.

[0595] Additionally, the server uses an emotion engine to recognize the user's emotional state. It analyzes the user's facial expressions and voice data to identify their current emotion (happiness, sadness, stress, etc.). Based on the recognized emotion, the visual content is optimized. For example, if the user is feeling stressed, the visual color tone will be changed to a more relaxing one.

[0596] The customized and optimized visual content is then sent from the server to the user's terminal or visual device, such as an augmented reality (AR) glass or a virtual reality (VR) headset, which then displays the received visual content in real time, allowing the user to instantly visualize the information.

[0597] For example, when a user with dyslexia tries to understand product information on an online shopping website, the process is as follows: The user uses their smartphone camera to capture the product description. The device uses OCR technology to extract text from the image and sends it to the server. The server then analyzes the text and converts it into a visually understandable format. This converted information is customized based on the user's profile information and further optimized for the user's emotional state. The customized visual content is then sent to the user's device, where it can be transparently understood.

[0598] Examples of prompts:

[0599] "Please take a picture of this product description to make it easier to understand visually."

[0600] In this way, the system of the present invention can provide textual information in a visually easy-to-understand format when a user with dyslexia is shopping online, and can also provide visual content optimized according to the user's emotional state.

[0601] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0602] Step 1:

[0603] A user uses a smartphone or tablet to capture text information to read. Specifically, they open a camera app and take a picture of a product description or review. The input is the image taken with the smartphone camera, and the output is the captured image data.

[0604] Step 2:

[0605] The device extracts text information from the acquired image data using OCR technology. In this process, the image data is input, and an OCR engine (e.g., Tesseract) is used to recognize characters, resulting in text data as the output.

[0606] Step 3:

[0607] Text data is sent from a terminal to a server. The input is text data that is transferred to the server via a communication network. The output is the text data that arrives at the server.

[0608] Step 4:

[0609] The server analyzes the received text data and converts it into a form that is easy to understand visually. Specifically, it performs semantic analysis using a natural language processing model (e.g., BERT or GPT) and generates images and videos using a generative AI model (e.g., GAN). The input is text data and the output is visual content (images and videos).

[0610] Step 5:

[0611] The server customizes the visual content based on the user's profile information, for example by adjusting colors to accommodate color blindness or changing text size. The inputs to this process are the user's profile information and the visual content, and the output is the customized visual content.

[0612] Step 6:

[0613] The server uses an emotion engine to recognize the user's emotional state. Specifically, it analyzes facial expression images and voice data provided by the user to identify emotions. The input of this process is emotion data (facial expression images and voice data), and the output is the recognized emotional state.

[0614] Step 7:

[0615] The server further optimizes the visual content based on the recognized emotional state, for example, changing the color tones to a more relaxing color if the user is stressed. The inputs to this process are the emotional state and the customized visual content, and the output is the optimized visual content.

[0616] Step 8:

[0617] The optimized visual content is sent from the server to the user's device or visual device. The input is the optimized visual content, which is sent to the user's device or AR glasses via a communication network. The output is the visual content displayed on the user's device or visual device.

[0618] Step 9:

[0619] The user checks the visual content displayed on the terminal or visual device and visually understands the required information. The input is the optimized content displayed on the visual device, and the output is the user's understanding. This step allows the user to visually understand the information comfortably.

[0620] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0621] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0622] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0623] [Third embodiment]

[0624] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0625] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0626] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0627] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0628] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0629] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0630] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0631] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0632] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0633] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0634] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0635] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0636] The system of the present invention is designed to help users with dyslexia visually understand textual information in their daily lives and studies. This system involves a series of processes in which the user's terminal acquires textual information, which is then converted into visual content by a server, which customizes the content, and then transmits it to a visual device.

[0637] System configuration and program handling

[0638] 1. Obtaining text information

[0639] Users use devices such as smartphones or AR glasses to capture the text information they want to read. Specifically, they can take a picture of a sign or a textbook page with a camera. The device then extracts the text information from the captured image using OCR (optical character recognition) technology. The extracted text information is then sent to a server via a communications network.

[0640] 2. Analyzing textual information and converting it into visual content

[0641] The server receives the text information sent from the device and performs syntactic analysis. Syntactic analysis allows the server to understand the meaning and structure of the text information. The server then uses a natural language processing (NLP) model to convert the text information into a visually understandable format. This converted text information is then further converted into images or videos using generative AI (e.g., GAN: Generative Artificial Intelligence Network).

[0642] 3. Individual customization

[0643] The server customizes the generated visual content based on the user's profile information (type of visual impairment, customization preferences, etc.) For example, it may emphasize certain colors for color-blind users, or adjust font size and placement to make them easier to understand. The customized visual content is temporarily stored in a database on the server.

[0644] 4. Transfer and display on a visual device

[0645] The server sends the customized visual content to the user's device or visual device (e.g., AR glasses or VR headset). The device displays the received visual content, allowing the user to visually confirm the text information in real time. This display makes it easy for users with disabilities, such as dyslexia, to understand the information.

[0646] Specific examples

[0647] When reading signs in the city

[0648] A user points their smartphone camera at a sign and captures it. The device extracts text from the sign image and sends it to the server. The server then analyzes the received text information and uses generative AI to convert the text into visual content. The server then customizes the visual content based on the user's profile information and sends it to the user's AR device. Finally, the user can visually understand the information on the sign through the AR device.

[0649] When reading the contents of a textbook

[0650] A student user takes a photo of a textbook page with their smartphone. The device extracts text from the textbook image and sends it to the server. The server then analyzes the received text information and converts it into a visually understandable format. The server then customizes it based on the user's profile information and sends the customized visual content to the user's device. The user can then view the textbook content in an easy-to-understand format through their visual device.

[0651] In this way, the system of the present invention provides a new method for making textual information easier to understand visually for users with dyslexia.

[0652] The processing flow will be explained below.

[0653] Step 1:

[0654] A user uses a device such as a smartphone or AR glasses to capture text information to be read. For example, the user takes a picture of a sign or a page in a textbook with a camera. The device then acquires the captured image data.

[0655] Step 2:

[0656] The device extracts text information from the acquired image data using OCR (Optical Character Recognition) technology, which converts the characters in the image into digital text. The device then sends the extracted text information to the server.

[0657] Step 3:

[0658] The server receives the text information sent from the terminal, analyzes the received text information, and performs syntax analysis. This analysis allows the meaning and structure of the text to be understood.

[0659] Step 4:

[0660] The server uses a natural language processing (NLP) model to convert the parsed text information into a visually understandable format, and then invokes a generative AI (e.g., GAN: Generative Artificial Intelligence Network) to convert the text into visual content such as images or videos, thereby generating the text information in a visual format.

[0661] Step 5:

[0662] The server references the user's profile information, which includes the user's type of visual impairment and customization preferences, and customizes the generated visual content based on this information. For example, it may enhance certain colors for users with color blindness, or change font size or layout.

[0663] Step 6:

[0664] The server then sends the customized visual content to the user's device or visual device (such as AR glasses or VR headset), in real time, allowing the user to view the information instantly.

[0665] Step 7:

[0666] The device displays the received visual content. By looking at the visual content generated through the device or visual device, the user can visually understand the text information they have read. This makes it easier for people with disabilities, such as dyslexia, to understand the information.

[0667] Example 1

[0668] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0669] Users with dyslexia have difficulty visually understanding textual information in their daily lives and in their studies. This problem is particularly pronounced when reading books or recognizing street signs. Conventional methods lack an efficient system for converting textual information into a visually understandable format. Given this background, there is a need to provide a method that enables users with dyslexia to more easily visually understand textual information.

[0670] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0671] In this invention, the server includes means for parsing text information and converting it into visual content, means for converting it into a visually understandable format using a generative AI model, and means for customizing the visual content based on user profile information, thereby enabling users with dyslexia to convert and display text information in a visually understandable format in real time.

[0672] "Text information" refers to information expressed as characters or sentences.

[0673] A "communications network" is an infrastructure for transmitting and receiving data.

[0674] A "server" is a computer system that processes and stores data on a network.

[0675] "Syntax analysis" is the process of understanding the grammar and structure of given text data.

[0676] "Visual content" refers to media that is displayed visually, such as images and videos.

[0677] A "generative AI model" is a learning algorithm for generating data using artificial intelligence techniques.

[0678] "User profile information" means information about an individual user, such as settings, preferences, and characteristics.

[0679] "Customization" refers to changing settings and content to suit individual needs and preferences.

[0680] A "visual device" is a device for displaying visual information, including AR devices and VR devices.

[0681] "Optical character recognition (OCR) technology" is a technology that extracts character information from image data.

[0682] This invention is a system for supporting users with dyslexia to visually understand textual information in their daily lives and studies. This system involves a series of processes: a terminal acquires textual information, a server converts it into visual content, customizes it, and sends it to a visual device.

[0683] First, the user captures the text information they want to read using a device such as a smartphone or AR glasses. Specifically, the user uses the smartphone camera to take a picture of a textbook page or a sign in the city. The device then uses OCR (optical character recognition) technology to extract the text information from the captured image. This extraction can be done using, for example, the Google Cloud Vision API.

[0684] The device then sends the extracted text information to a server over a communications network, typically using an HTTP POST request, with an endpoint set up using, for example, AWS API Gateway.

[0685] The server first parses the received text information. For example, the server uses the Python library spaCy to analyze the structure and meaning of the text. The server then uses an NLP (natural language processing) model to convert the text information into a visually understandable format. Specifically, OpenAI's generative AI model (e.g., GPT-3) or a generative artificial network (GAN) model is used.

[0686] Additionally, the server customizes the generated visual content based on the user's profile information, for example, increasing font size or enhancing colors for a user with dyslexia, based on the user's preference database.

[0687] Finally, the server sends the customized visual content to the user's device or visual device (e.g., AR glasses or VR headset), again using an HTTP POST request. The device displays the received visual content, allowing the user to visually confirm text information in real time. This helps users with dyslexia to more easily understand text information in their daily lives.

[0688] Specific examples

[0689] When reading signs in the city

[0690] The user points their smartphone camera at a sign and captures it. The device uses the Google Cloud Vision API to extract text from the sign image and sends it to the server via AWS API Gateway. The server then analyzes the received text using spaCy and converts it into a visually understandable format using a generative AI model. The text is then further customized based on the user's profile information, and the final visual content is sent to the user's AR device. Through the AR device, the user can visually perceive the sign's information, for example, with larger, more emphasized text.

[0691] When reading the contents of a textbook

[0692] A student user takes a photo of a textbook page with their smartphone camera. The device uses the Google Cloud Vision API to extract text information from the textbook image and sends it to a server using AWS API Gateway. The server analyzes the received text information and converts it into visual content that is easy to understand using OpenAI's GPT-3 or GAN. This content is customized based on the user's profile information and then sent to the vision device. Finally, the user can view the textbook content on the vision device with text alignment and font size adjusted.

[0693] Prompt Sentence Examples

[0694] An example of a prompt in a generative AI model is, "Please convert this text information into a visually easy-to-understand image. User profile information is as follows: type of visual impairment is color blindness, customization preference is to set text background color to blue."

[0695] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0696] Step 1:

[0697] Users capture text information using devices such as smartphones or AR glasses. For example, they take a picture of a page in a textbook or a sign in the city using the smartphone camera.

[0698] Input: Physical text information (textbook pages, signs, etc.)

[0699] Output: Captured image data

[0700] Step 2:

[0701] The device uses OCR technology on the captured image to extract text information, using the Google Cloud Vision API, and then converts the extracted text information into a data format.

[0702] Input: Photographed image data

[0703] Output: Extracted text information data

[0704] Step 3:

[0705] The device sends the extracted text information to a server via a communication network. The device securely transfers the data to the server using an HTTP POST request, for example, using AWS API Gateway.

[0706] Input: Extracted text information data

[0707] Output: Text information sent to the server

[0708] Step 4:

[0709] The server receives the text information sent from the terminal and performs syntax analysis. The server uses a Python library (e.g., spaCy) to analyze the structure of the text and understand its grammar and meaning.

[0710] Input: Text information sent to the server

[0711] Output: Structure data of parsed text information

[0712] Step 5:

[0713] The server uses natural language processing (NLP) models to convert text information into a visually understandable format. The server then uses generative AI models (e.g., GPT-3) or GANs to convert the analyzed text information into visual content (images and videos).

[0714] Input: Structure data of parsed text information

[0715] Output: Visual content that is easy to understand

[0716] Step 6:

[0717] The server customizes the generated visual content based on the user's profile information, such as increasing font size or highlighting certain colors for a user with dyslexia, based on the user's preference database.

[0718] Input: User profile information, visual content

[0719] Output: Customized visual content

[0720] Step 7:

[0721] The server sends the customized visual content to the user's device or viewing device, again using an HTTP POST request to transfer the data. The device then displays the received visual content.

[0722] Input: Customized visual content

[0723] Output: Visual content displayed on the device

[0724] Step 8:

[0725] Users can see the received visual content in real time, making it easier for users with disabilities such as dyslexia to understand the information.

[0726] Input: Visual content displayed on the device

[0727] Output: Visually understood text information

[0728] (Application example 1)

[0729] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0730] Users with dyslexia have difficulty visually understanding written information such as product labels and guide signs in physical stores. This makes it difficult for users to obtain appropriate information, limiting their ability to select products and use store facilities. There is a need for a system that can support such users and enable them to shop and use physical stores comfortably.

[0731] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0732] In this invention, the server includes means for capturing product labels and guide signs in a physical store using the smart glasses and converting the text information into a visually understandable format, means for customizing the images or videos based on the user's profile information, and means for displaying the generated visual content on the smart glasses so that the user can view the visual information in real time, thereby enabling users with dyslexia to easily understand products and guide signs in a physical store and enjoy shopping and facility use.

[0733] "Text information" refers to character data found in books, signs, product labels, and the like.

[0734] "Image or video" refers to still images or dynamic video data that visually represent text information.

[0735] "User profile information" refers to personalized information such as the type of visual impairment the user has and their customization preferences.

[0736] A "visual device" is a device that allows a user to receive visual information, specifically smart glasses, augmented reality (AR) devices, and virtual reality (VR) devices.

[0737] A "brick and mortar store" is a physical location for selling goods or providing services.

[0738] "Capture" refers to the act of using a device such as a camera to obtain images of product labels or guide signs in a physical store.

[0739] "Optical character recognition (OCR) technology" is a technology for automatically extracting text information from images.

[0740] "Smart glasses" are eyeglass-type devices equipped with camera and display functions, allowing users to check visual information in real time.

[0741] This invention is a system for supporting users with dyslexia to visually understand text information such as product labels and guide signs in physical stores. A specific embodiment of this system is described below.

[0742] First, a user wears smart glasses and captures product labels and guide signs while walking around a physical store. The smart glasses have a camera function and extract text information from the captured images using optical character recognition (OCR) technology (e.g., Tesseract OCR). The extracted text information is then sent from the smart glasses to a server via a communication network.

[0743] The server then parses the received text information and uses a natural language processing (NLP) model (e.g., BERT, GPT-3) to understand its meaning and structure. The text information is converted into a visually understandable format, which is then transformed into visual content in the form of images or videos using a generative AI model (e.g., GAN, DALL-E).

[0744] The server then customizes the generated visual content based on the user's profile information (e.g., type of visual impairment, customization preferences, etc.), for example, by highlighting certain colors or adjusting font size. The customized visual content is temporarily stored in a database within the server.

[0745] Finally, the server transmits the customized visual content to the smart glasses via a communication network, and the smart glasses display the received visual content for the user to view in real time.

[0746] As a concrete example, consider a user with dyslexia capturing a product label in the detergent section of a brick-and-mortar store. The user uses the camera in their smart glasses to capture the label: "Detergent - 750 yen special offer!"

[0747] 1. The camera takes a picture of the label and uses OCR technology (e.g., Tesseract OCR) to extract the text information: "Detergent - Special Price: 750 yen!"

[0748] 2. This text information is sent from the smart glasses to the server.

[0749] 3. The server uses an NLP model (e.g., GPT-3) to convert the text information into a visually understandable format, and a generative AI model (e.g., GAN) to generate a customized image (e.g., highlighting the "On Sale!" part in red).

[0750] 4. The server sends the generated visual content to the smart glasses, and the user can view the received visual content in real time.

[0751] A specific example of a prompt is "Please convert the following text into a visually understandable format. Text: 'Detergent - Special Offer for 750 yen!'" Based on this prompt, the server converts the text information using an NLP model and then generates visual content that is easy to understand using a generative AI model.

[0752] This allows users with dyslexia to easily understand products and information in physical stores, allowing them to shop and use facilities comfortably.

[0753] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0754] Step 1:

[0755] The user wears smart glasses and uses the camera to capture product labels and information signs in a physical store. The input is image data captured by the smart glasses' camera. The output is saved as image data on the device. Specifically, the user points the camera at the information they want to see and presses a button to capture the image.

[0756] Step 2:

[0757] The device extracts text information from the captured image using OCR (Optical Character Recognition) technology (e.g., Tesseract OCR). The input is the captured image data. The output is the extracted text information. Specifically, the OCR software analyzes the image, recognizes the text, and extracts it as character data.

[0758] Step 3:

[0759] The extracted text information is sent from the device to the server via a communication network. The input is the text information extracted by the OCR. The output is the text information received by the server. Specifically, the device uploads the text information to the server via an Internet connection.

[0760] Step 4:

[0761] The server uses an NLP model (e.g., BERT, GPT-3) to parse the received text information. The input is the text information sent from the device. The output is the text information converted to make it easier to understand visually. Specifically, the NLP model analyzes the meaning and structure of the text and performs natural language processing.

[0762] Step 5:

[0763] The server uses a generative AI model (e.g., GAN, DALL-E) to generate images and videos based on the information obtained using the NLP model and converts it into a visually easy-to-understand format. The input is text information analyzed by the NLP model. The output is images or videos in a visually easy-to-understand format. Specifically, the generative AI model creates appropriate visual content.

[0764] Step 6:

[0765] The server customizes the generated visual content based on the user's profile information. The input is images or videos converted to be visually easy to understand, along with the user's profile information. The output is customized visual content. Specific actions include customization to suit the user's needs, such as enhancing colors or adjusting font sizes.

[0766] Step 7:

[0767] The server transmits the customized visual content to the user's smart glasses via a communication network. The input is the customized visual content. The output is the visual content received by the smart glasses. Specific operations include data transfer from the server to the smart glasses.

[0768] Step 8:

[0769] The terminal displays the received visual content, and the user can view it in real time. The input is the customized visual content sent from the server. The output is the information that the user visually views. In concrete terms, the visual content is displayed on the display of the smart glasses, allowing the user to view the information.

[0770] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0771] The system of the present invention supports users with dyslexia in visually understanding textual information in their daily lives and studies. This embodiment also adds a function to recognize the user's emotions and optimize visual content based on those emotions. This system involves a series of processes: the user's device acquires textual information, the server converts it into customized visual content, and the server further optimizes it based on the user's emotions and transmits it to the visual device.

[0772] System configuration and program handling

[0773] 1. Obtaining text information

[0774] A user uses a device such as a smartphone or AR glasses to capture text information to be read. For example, they take a picture of a sign or a page in a textbook with a camera. The device then extracts the text information from the captured image using OCR (optical character recognition) technology. The extracted text information is then sent to a server via a communication network.

[0775] 2. Analyzing textual information and converting it into visual content

[0776] The server receives the text information sent from the device and performs syntactic analysis. Syntactic analysis understands the meaning and structure of the text. The server then uses a natural language processing (NLP) model to convert the text information into a visually understandable format. This converted text information is then further converted into images or videos using generative AI (e.g., GAN: Generative Artificial Intelligence Network).

[0777] 3. Individual customization

[0778] The server customizes the generated visual content based on the user's profile information (type of visual impairment, customization preferences, etc.) For example, it may emphasize certain colors for color-blind users, or adjust font size and placement to make them easier to understand. The customized visual content is temporarily stored in a database on the server.

[0779] 4. Optimization by Emotion Engine

[0780] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's facial expression data or voice data to identify the emotion the user is currently feeling (e.g., joy, sadness, surprise, anger, etc.). Based on this emotion data, the server further adjusts the customization of the visual content. For example, if the user is feeling stressed, the server may change the color tone to one that is visually more relaxing.

[0781] 5. Transfer and display on a visual device

[0782] The server sends the customized visual content to the user's device or visual device (e.g., AR glasses or VR headset). This transfer occurs in real time, allowing the user to instantly view the information. The device then displays the received visual content, allowing the user to visually interpret the text information.

[0783] Specific examples

[0784] When reading signs in the city

[0785] The user points their smartphone camera at a sign and captures it. The device extracts text from the sign image and sends it to the server. The server analyzes the received text information and uses generative AI to convert the text into visual content. The server then customizes the content based on the user's profile information and uses an emotion engine to recognize and optimize the user's emotional state. The customized visual content is then sent to the device, allowing the user to visually understand information about the city.

[0786] When reading the contents of a textbook

[0787] A student user takes a photo of a textbook page with their smartphone. The device extracts text from the textbook image and sends it to the server. The server analyzes the received text information and converts it into a visually understandable format. The server then customizes the content based on the user's profile information, and an emotion engine analyzes the user's concentration level and emotional state to optimize the visual content. The user can view the textbook content in an easy-to-understand format through the adjusted visual content.

[0788] In this way, the system of the present invention not only makes it easier for users with dyslexia to visually understand textual information, but also provides visual content that is optimally adjusted according to the user's emotional state.

[0789] The processing flow will be explained below.

[0790] Step 1:

[0791] A user uses a device such as a smartphone or AR glasses to capture text information to be read. For example, the user takes a picture of a sign or a page in a textbook with a camera. The device then acquires the captured image data.

[0792] Step 2:

[0793] The device extracts text information from the acquired image data using OCR (Optical Character Recognition) technology, which converts the characters in the image into digital text. The device then sends the extracted text information to the server.

[0794] Step 3:

[0795] The server receives the text information sent from the terminal, analyzes the received text information, and performs syntax analysis. This analysis allows the meaning and structure of the text to be understood.

[0796] Step 4:

[0797] The server uses a natural language processing (NLP) model to convert the parsed text information into a visually understandable format, and then invokes a generative AI (e.g., GAN: Generative Artificial Intelligence Network) to convert the text into visual content such as images or videos, thereby generating the text information in a visual format.

[0798] Step 5:

[0799] The server references the user's profile information, which includes the user's type of visual impairment and customization preferences, and customizes the generated visual content based on this information. For example, it may enhance certain colors for users with color blindness, or change font size or layout.

[0800] Step 6:

[0801] The server recognizes the user's emotions using an emotion engine, which analyzes the user's facial expression data or voice data to identify the emotion the user is currently feeling (e.g., joy, sadness, surprise, anger, etc.).

[0802] Step 7:

[0803] The server further adjusts the visual content based on the user's emotional state, for example, changing the color tones to be more visually relaxing if the user is feeling stressed.

[0804] Step 8:

[0805] The server then sends the customized visual content to the user's device or viewing device (e.g., AR glasses or VR headset). This transfer occurs in real time, allowing the user to view the information instantly.

[0806] Step 9:

[0807] The device displays the received visual content. By looking at the visual content generated through the device or visual device, the user can visually understand the text information they have read. This makes it easier for people with disabilities, such as dyslexia, to understand the information.

[0808] Example 2

[0809] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0810] Conventional visual content generation systems focus on converting text information into something visually easy to understand, but do not fully consider the visual characteristics and emotional state of each user. This makes it difficult to provide optimal visual information, especially for users with dyslexia. Furthermore, they lack flexible customization to allow users to obtain information without feeling stressed or burdened.

[0811] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring text information, means for converting the text information into images or videos, means for customizing the images or videos based on the user's profile information, means for analyzing the user's emotional state and optimizing the visual content, and means for transmitting the customized and optimized images or videos to the user's visual device. This makes it possible to provide optimized visual content that takes into account the visual characteristics and emotional state of each user.

[0812] "Text information" is data of characters and sentences that the user should read.

[0813] The "means for acquiring" is a mechanism for acquiring text information using the user's device.

[0814] The "means for converting into images or videos" is a mechanism for converting acquired text information into images or videos that are visually easy to understand.

[0815] "User profile information" is data about an individual user's visual characteristics, preferences, and disabilities.

[0816] A "means for customizing" is a mechanism for optimizing and tailoring generated visual content to a particular user based on the user's profile information.

[0817] The "emotional state of the user" refers to the mood or state of mind of the user that can be understood through detected facial expressions and voice data.

[0818] The "analyzing means" is a mechanism for analyzing facial expression data and voice data to understand the emotional state of the user.

[0819] The "optimizing means" is a mechanism that further adjusts the visual content based on the user's emotional state to improve user comfort.

[0820] A "visual device" is a device used by a user to view visual content, including, for example, an augmented reality device or a virtual reality device.

[0821] A "transmitting means" is a mechanism that transfers customized and optimized visual content to a user's visual device.

[0822] A "system" is a collection of devices and software that include each of these means and operate in conjunction with one another.

[0823] The system of the present invention is designed to help users with dyslexia visually understand written information in their daily lives and studies. This system involves a series of processes: acquiring the user's text information, analyzing and converting it on the server, and further customizing and optimizing it based on the user's profile information and emotional state.

[0824] Specific examples of hardware and software used

[0825] Hardware:

[0826] Smartphone

[0827] Augmented Reality (AR) Devices

[0828] Virtual Reality (VR) Devices

[0829] software:

[0830] OCR (optical character recognition) technology

[0831] Parsing Engine

[0832] Natural Language Processing (NLP) Models

[0833] Generative AI (e.g., GAN: Generative Adversarial Networks)

[0834] Sentiment Analysis Engine

[0835] System operation example

[0836] When reading signs in the city

[0837] The user points their smartphone camera at a sign and captures it. The device uses OCR technology to extract text information, which is then sent to the server. The server then parses the received text and uses generative AI to convert it into visually understandable images or videos. The server then customizes the visual content based on the user's profile information and optimizes it using an emotion engine to recognize the user's emotional state. The customized and optimized visual content is then sent to the device, allowing the user to instantly visually understand the information on signs around town.

[0838] When reading the contents of a textbook

[0839] A student user takes a photo of a textbook page with their smartphone. The device uses OCR technology to extract text from the textbook image and sends it to the server. The server then parses the received text information and uses generative AI to convert it into a visually understandable format. The server then customizes the visual content based on the user's profile information and optimizes it using an emotion engine to analyze the user's concentration level and emotional state. The user can visually understand the content of the textbook through the tailored visual content.

[0840] Prompt Sentence Examples

[0841] "Use GANs to run a program that translates text information into visual content for dyslexic users, and optimize it for the user's emotional state. For example, if the user is stressed, adjust the color tones to be more relaxing."

[0842] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0843] Step 1: Obtaining text information

[0844] A user uses a device such as a smartphone or AR glasses to capture text information using the camera. For example, the user takes a picture of a sign in the city or a page in a textbook.

[0845] Input: Image data taken by the user

[0846] The device inputs the image data into OCR (optical character recognition) software to extract text information.

[0847] Output: Extracted text information

[0848] The terminal compresses the extracted text information, performs an error check, encrypts it, and transmits it to the server.

[0849] Step 2: Parsing and transforming text information

[0850] The server receives the text information sent from the terminal and stores it in a database.

[0851] Input: Text information sent from the device

[0852] The server uses natural language processing (NLP) models to analyze the text and understand its meaning and context, including morphological analysis and dependency analysis.

[0853] Output: Parsed text information

[0854] The server uses generative AI (e.g., GAN) based on the analyzed text information to convert it into visually easy-to-understand image or video content.

[0855] Output: The generated visual content (images or videos)

[0856] Step 3: Individual customization

[0857] The server refers to the user's profile information (visual characteristics, preferences, etc.) stored in a database.

[0858] Input: Generated visual content, user profile information

[0859] The server customizes the visual content based on the user's profile information, for example by enhancing certain colors for color-blind users or adjusting font size and placement.

[0860] Output: Customized visual content

[0861] The server temporarily stores the customized visual content in a database.

[0862] Step 4: Optimizing with an Emotional Engine

[0863] The server inputs facial expression data and voice data acquired from the terminal into an emotion engine and analyzes the user's emotional state.

[0864] Input: facial expression data, voice data

[0865] The server uses an emotion engine to identify the user's emotion (joy, sadness, surprise, anger, etc.).

[0866] Output: Emotion analysis results

[0867] Based on the results of the emotion analysis, the server adjusts the color tone and layout of the visual content, optimizing it so that users can view information more comfortably.

[0868] Output: Optimized visual content

[0869] Step 5: Transfer and display to a visual device

[0870] The server transmits optimized visual content in real time to the user's device or visual device (such as AR glasses or VR headset).

[0871] Input: Optimized visual content

[0872] The terminal decodes the received visual content and displays it on its display.

[0873] Output: The displayed visual content

[0874] Users view customized and optimized visual content through their terminals and visual devices and visually understand textual information.

[0875] (Application example 2)

[0876] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0877] Traditionally, users with dyslexia have had difficulty visually understanding textual information. This can also hinder their ability to understand textual information during everyday activities such as online shopping and information searches, resulting in stress and misunderstandings. Furthermore, visual content provided without considering the user's emotional state can make comprehension even more difficult. The present invention aims to solve these problems and enable users with dyslexia to easily understand textual information and live their daily lives comfortably.

[0878] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring text information, means for analyzing the acquired text information and converting it into a form that is visually easy to understand, means for recognizing a user's emotion and optimizing visual content based on the emotion, means for customizing the visual content based on profile information, and means for transmitting the customized and optimized visual content to the user's terminal or visual device. This makes it easier for users with dyslexia to visually understand text information, and further enables the provision of optimal content according to their emotional state.

[0879] "Text information" is data of characters or sentences that are visually displayed.

[0880] "Means for acquiring text information" refers to a method of capturing visual characters or sentences as electronic data using a device such as a camera or scanner.

[0881] "Analysis" refers to processing the acquired text information to understand its meaning and structure.

[0882] "Means of converting into a visually easy-to-understand form" refers to methods of converting text information into a visually easy-to-understand form such as an image or video.

[0883] "Means for recognizing a user's emotions and optimizing visual content based on those emotions" refers to a method for analyzing a user's current emotional state from facial expressions and voice, and adjusting visual content to appropriately reflect those emotions.

[0884] "Profile information" refers to information such as the type of visual impairment and customization preferences of each individual user.

[0885] "Means for customizing visual content" refers to methods for adjusting color settings, font size, placement, etc. to suit each user based on profile information.

[0886] "Visual devices" are devices for visually displaying digital information, such as AR glasses and VR headsets.

[0887] A "server" is a centralized computer system for processing and managing data.

[0888] A "user's terminal" is a device that a user can operate at hand, such as a smartphone or tablet.

[0889] A "natural language processing model" is a learning model for analyzing and generating text data.

[0890] An "emotion engine" is software for analyzing a user's emotional state.

[0891] A "generative AI model" is an artificial intelligence model that uses neural networks to generate new data.

[0892] This invention provides a system that helps users with dyslexia visually understand text information in daily life and online shopping. In this embodiment, this system is realized through the following steps.

[0893] The server analyzes the text information sent by the user and converts it into a visually understandable format. First, the user captures the text information they want to obtain using a device such as a smartphone or tablet. From this captured image, the text information is extracted using optical character recognition (OCR) technology. The extracted text information is then sent to the server.

[0894] The server performs syntactic analysis on the received text to understand its meaning and structure, then uses a natural language processing (NLP) model to convert the text into a visually understandable format, which is then converted into an image or video using a generative AI model (e.g., a generative neural network).

[0895] The server then customizes the visual content based on the profile information, which may include, for example, the type of visual impairment and customization preferences, such as color enhancement or font size adjustment.

[0896] Additionally, the server uses an emotion engine to recognize the user's emotional state. It analyzes the user's facial expressions and voice data to identify their current emotion (happiness, sadness, stress, etc.). Based on the recognized emotion, the visual content is optimized. For example, if the user is feeling stressed, the visual color tone will be changed to a more relaxing one.

[0897] The customized and optimized visual content is then sent from the server to the user's terminal or visual device, such as an augmented reality (AR) glass or a virtual reality (VR) headset, which then displays the received visual content in real time, allowing the user to instantly visualize the information.

[0898] For example, when a user with dyslexia tries to understand product information on an online shopping website, the process is as follows: The user uses their smartphone camera to capture the product description. The device uses OCR technology to extract text from the image and sends it to the server. The server then analyzes the text and converts it into a visually understandable format. This converted information is customized based on the user's profile information and further optimized for the user's emotional state. The customized visual content is then sent to the user's device, where it can be transparently understood.

[0899] Examples of prompts:

[0900] "Please take a picture of this product description to make it easier to understand visually."

[0901] In this way, the system of the present invention can provide textual information in a visually easy-to-understand format when a user with dyslexia is shopping online, and can also provide visual content optimized according to the user's emotional state.

[0902] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0903] Step 1:

[0904] A user uses a smartphone or tablet to capture text information to read. Specifically, they open a camera app and take a picture of a product description or review. The input is the image taken with the smartphone camera, and the output is the captured image data.

[0905] Step 2:

[0906] The device extracts text information from the acquired image data using OCR technology. In this process, the image data is input, and an OCR engine (e.g., Tesseract) is used to recognize characters, resulting in text data as the output.

[0907] Step 3:

[0908] Text data is sent from a terminal to a server. The input is text data that is transferred to the server via a communication network. The output is the text data that arrives at the server.

[0909] Step 4:

[0910] The server analyzes the received text data and converts it into a form that is easy to understand visually. Specifically, it performs semantic analysis using a natural language processing model (e.g., BERT or GPT) and generates images and videos using a generative AI model (e.g., GAN). The input is text data and the output is visual content (images and videos).

[0911] Step 5:

[0912] The server customizes the visual content based on the user's profile information, for example by adjusting colors to accommodate color blindness or changing text size. The inputs to this process are the user's profile information and the visual content, and the output is the customized visual content.

[0913] Step 6:

[0914] The server uses an emotion engine to recognize the user's emotional state. Specifically, it analyzes facial expression images and voice data provided by the user to identify emotions. The input of this process is emotion data (facial expression images and voice data), and the output is the recognized emotional state.

[0915] Step 7:

[0916] The server further optimizes the visual content based on the recognized emotional state, for example, changing the color tones to a more relaxing color if the user is stressed. The inputs to this process are the emotional state and the customized visual content, and the output is the optimized visual content.

[0917] Step 8:

[0918] The optimized visual content is sent from the server to the user's device or visual device. The input is the optimized visual content, which is sent to the user's device or AR glasses via a communication network. The output is the visual content displayed on the user's device or visual device.

[0919] Step 9:

[0920] The user checks the visual content displayed on the terminal or visual device and visually understands the required information. The input is the optimized content displayed on the visual device, and the output is the user's understanding. This step allows the user to visually understand the information comfortably.

[0921] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0922] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0923] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0924] [Fourth embodiment]

[0925] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0926] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0927] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0928] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0929] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0930] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0931] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0932] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0933] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0934] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0935] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0936] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0937] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0938] The system of the present invention is designed to help users with dyslexia visually understand textual information in their daily lives and studies. This system involves a series of processes in which the user's terminal acquires textual information, which is then converted into visual content by a server, which customizes the content, and then transmits it to a visual device.

[0939] System configuration and program handling

[0940] 1. Obtaining text information

[0941] Users use devices such as smartphones or AR glasses to capture the text information they want to read. Specifically, they can take a picture of a sign or a textbook page with a camera. The device then extracts the text information from the captured image using OCR (optical character recognition) technology. The extracted text information is then sent to a server via a communications network.

[0942] 2. Analyzing textual information and converting it into visual content

[0943] The server receives the text information sent from the device and performs syntactic analysis. Syntactic analysis allows the server to understand the meaning and structure of the text information. The server then uses a natural language processing (NLP) model to convert the text information into a visually understandable format. This converted text information is then further converted into images or videos using generative AI (e.g., GAN: Generative Artificial Intelligence Network).

[0944] 3. Individual customization

[0945] The server customizes the generated visual content based on the user's profile information (type of visual impairment, customization preferences, etc.) For example, it may emphasize certain colors for color-blind users, or adjust font size and placement to make them easier to understand. The customized visual content is temporarily stored in a database on the server.

[0946] 4. Transfer and display on a visual device

[0947] The server sends the customized visual content to the user's device or visual device (e.g., AR glasses or VR headset). The device displays the received visual content, allowing the user to visually confirm the text information in real time. This display makes it easy for users with disabilities, such as dyslexia, to understand the information.

[0948] Specific examples

[0949] When reading signs in the city

[0950] A user points their smartphone camera at a sign and captures it. The device extracts text from the sign image and sends it to the server. The server then analyzes the received text information and uses generative AI to convert the text into visual content. The server then customizes the visual content based on the user's profile information and sends it to the user's AR device. Finally, the user can visually understand the information on the sign through the AR device.

[0951] When reading the contents of a textbook

[0952] A student user takes a photo of a textbook page with their smartphone. The device extracts text from the textbook image and sends it to the server. The server then analyzes the received text information and converts it into a visually understandable format. The server then customizes it based on the user's profile information and sends the customized visual content to the user's device. The user can then view the textbook content in an easy-to-understand format through their visual device.

[0953] In this way, the system of the present invention provides a new method for making textual information easier to understand visually for users with dyslexia.

[0954] The processing flow will be explained below.

[0955] Step 1:

[0956] A user uses a device such as a smartphone or AR glasses to capture text information to be read. For example, the user takes a picture of a sign or a page in a textbook with a camera. The device then acquires the captured image data.

[0957] Step 2:

[0958] The device extracts text information from the acquired image data using OCR (Optical Character Recognition) technology, which converts the characters in the image into digital text. The device then sends the extracted text information to the server.

[0959] Step 3:

[0960] The server receives the text information sent from the terminal, analyzes the received text information, and performs syntax analysis. This analysis allows the meaning and structure of the text to be understood.

[0961] Step 4:

[0962] The server uses a natural language processing (NLP) model to convert the parsed text information into a visually understandable format, and then invokes a generative AI (e.g., GAN: Generative Artificial Intelligence Network) to convert the text into visual content such as images or videos, thereby generating the text information in a visual format.

[0963] Step 5:

[0964] The server references the user's profile information, which includes the user's type of visual impairment and customization preferences, and customizes the generated visual content based on this information. For example, it may enhance certain colors for users with color blindness, or change font size or layout.

[0965] Step 6:

[0966] The server then sends the customized visual content to the user's device or visual device (such as AR glasses or VR headset), in real time, allowing the user to view the information instantly.

[0967] Step 7:

[0968] The device displays the received visual content. By looking at the visual content generated through the device or visual device, the user can visually understand the text information they have read. This makes it easier for people with disabilities, such as dyslexia, to understand the information.

[0969] Example 1

[0970] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0971] Users with dyslexia have difficulty visually understanding textual information in their daily lives and in their studies. This problem is particularly pronounced when reading books or recognizing street signs. Conventional methods lack an efficient system for converting textual information into a visually understandable format. Given this background, there is a need to provide a method that enables users with dyslexia to more easily visually understand textual information.

[0972] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0973] In this invention, the server includes means for parsing text information and converting it into visual content, means for converting it into a visually understandable format using a generative AI model, and means for customizing the visual content based on user profile information, thereby enabling users with dyslexia to convert and display text information in a visually understandable format in real time.

[0974] "Text information" refers to information expressed as characters or sentences.

[0975] A "communications network" is an infrastructure for transmitting and receiving data.

[0976] A "server" is a computer system that processes and stores data on a network.

[0977] "Syntax analysis" is the process of understanding the grammar and structure of given text data.

[0978] "Visual content" refers to media that is displayed visually, such as images and videos.

[0979] A "generative AI model" is a learning algorithm for generating data using artificial intelligence techniques.

[0980] "User profile information" means information about an individual user, such as settings, preferences, and characteristics.

[0981] "Customization" refers to changing settings and content to suit individual needs and preferences.

[0982] A "visual device" is a device for displaying visual information, including AR devices and VR devices.

[0983] "Optical character recognition (OCR) technology" is a technology that extracts character information from image data.

[0984] This invention is a system for supporting users with dyslexia to visually understand textual information in their daily lives and studies. This system involves a series of processes: a terminal acquires textual information, a server converts it into visual content, customizes it, and sends it to a visual device.

[0985] First, the user captures the text information they want to read using a device such as a smartphone or AR glasses. Specifically, the user uses the smartphone camera to take a picture of a textbook page or a sign in the city. The device then uses OCR (optical character recognition) technology to extract the text information from the captured image. This extraction can be done using, for example, the Google Cloud Vision API.

[0986] The device then sends the extracted text information to a server over a communications network, typically using an HTTP POST request, with an endpoint set up using, for example, AWS API Gateway.

[0987] The server first parses the received text information. For example, the server uses the Python library spaCy to analyze the structure and meaning of the text. The server then uses an NLP (natural language processing) model to convert the text information into a visually understandable format. Specifically, OpenAI's generative AI model (e.g., GPT-3) or a generative artificial network (GAN) model is used.

[0988] Additionally, the server customizes the generated visual content based on the user's profile information, for example, increasing font size or enhancing colors for a user with dyslexia, based on the user's preference database.

[0989] Finally, the server sends the customized visual content to the user's device or visual device (e.g., AR glasses or VR headset), again using an HTTP POST request. The device displays the received visual content, allowing the user to visually confirm text information in real time. This helps users with dyslexia to more easily understand text information in their daily lives.

[0990] Specific examples

[0991] When reading signs in the city

[0992] The user points their smartphone camera at a sign and captures it. The device uses the Google Cloud Vision API to extract text from the sign image and sends it to the server via AWS API Gateway. The server then analyzes the received text using spaCy and converts it into a visually understandable format using a generative AI model. The text is then further customized based on the user's profile information, and the final visual content is sent to the user's AR device. Through the AR device, the user can visually perceive the sign's information, for example, with larger, more emphasized text.

[0993] When reading the contents of a textbook

[0994] A student user takes a photo of a textbook page with their smartphone camera. The device uses the Google Cloud Vision API to extract text information from the textbook image and sends it to a server using AWS API Gateway. The server analyzes the received text information and converts it into visual content that is easy to understand using OpenAI's GPT-3 or GAN. This content is customized based on the user's profile information and then sent to the vision device. Finally, the user can view the textbook content on the vision device with text alignment and font size adjusted.

[0995] Prompt Sentence Examples

[0996] An example of a prompt in a generative AI model is, "Please convert this text information into a visually easy-to-understand image. User profile information is as follows: type of visual impairment is color blindness, customization preference is to set text background color to blue."

[0997] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0998] Step 1:

[0999] Users capture text information using devices such as smartphones or AR glasses. For example, they take a picture of a page in a textbook or a sign in the city using the smartphone camera.

[1000] Input: Physical text information (textbook pages, signs, etc.)

[1001] Output: Captured image data

[1002] Step 2:

[1003] The device uses OCR technology on the captured image to extract text information, using the Google Cloud Vision API, and then converts the extracted text information into a data format.

[1004] Input: Photographed image data

[1005] Output: Extracted text information data

[1006] Step 3:

[1007] The device sends the extracted text information to a server via a communication network. The device securely transfers the data to the server using an HTTP POST request, for example, using AWS API Gateway.

[1008] Input: Extracted text information data

[1009] Output: Text information sent to the server

[1010] Step 4:

[1011] The server receives the text information sent from the terminal and performs syntax analysis. The server uses a Python library (e.g., spaCy) to analyze the structure of the text and understand its grammar and meaning.

[1012] Input: Text information sent to the server

[1013] Output: Structure data of parsed text information

[1014] Step 5:

[1015] The server uses natural language processing (NLP) models to convert text information into a visually understandable format. The server then uses generative AI models (e.g., GPT-3) or GANs to convert the analyzed text information into visual content (images and videos).

[1016] Input: Structure data of parsed text information

[1017] Output: Visual content that is easy to understand

[1018] Step 6:

[1019] The server customizes the generated visual content based on the user's profile information, such as increasing font size or highlighting certain colors for a user with dyslexia, based on the user's preference database.

[1020] Input: User profile information, visual content

[1021] Output: Customized visual content

[1022] Step 7:

[1023] The server sends the customized visual content to the user's device or viewing device, again using an HTTP POST request to transfer the data. The device then displays the received visual content.

[1024] Input: Customized visual content

[1025] Output: Visual content displayed on the device

[1026] Step 8:

[1027] Users can see the received visual content in real time, making it easier for users with disabilities such as dyslexia to understand the information.

[1028] Input: Visual content displayed on the device

[1029] Output: Visually understood text information

[1030] (Application example 1)

[1031] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1032] Users with dyslexia have difficulty visually understanding written information such as product labels and guide signs in physical stores. This makes it difficult for users to obtain appropriate information, limiting their ability to select products and use store facilities. There is a need for a system that can support such users and enable them to shop and use physical stores comfortably.

[1033] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1034] In this invention, the server includes means for capturing product labels and guide signs in a physical store using the smart glasses and converting the text information into a visually understandable format, means for customizing the images or videos based on the user's profile information, and means for displaying the generated visual content on the smart glasses so that the user can view the visual information in real time, thereby enabling users with dyslexia to easily understand products and guide signs in a physical store and enjoy shopping and facility use.

[1035] "Text information" refers to character data found in books, signs, product labels, and the like.

[1036] "Image or video" refers to still images or dynamic video data that visually represent text information.

[1037] "User profile information" refers to personalized information such as the type of visual impairment the user has and their customization preferences.

[1038] A "visual device" is a device that allows a user to receive visual information, specifically smart glasses, augmented reality (AR) devices, and virtual reality (VR) devices.

[1039] A "brick and mortar store" is a physical location for selling goods or providing services.

[1040] "Capture" refers to the act of using a device such as a camera to obtain images of product labels or guide signs in a physical store.

[1041] "Optical character recognition (OCR) technology" is a technology for automatically extracting text information from images.

[1042] "Smart glasses" are eyeglass-type devices equipped with camera and display functions, allowing users to check visual information in real time.

[1043] This invention is a system for supporting users with dyslexia to visually understand text information such as product labels and guide signs in physical stores. A specific embodiment of this system is described below.

[1044] First, a user wears smart glasses and captures product labels and guide signs while walking around a physical store. The smart glasses have a camera function and extract text information from the captured images using optical character recognition (OCR) technology (e.g., Tesseract OCR). The extracted text information is then sent from the smart glasses to a server via a communication network.

[1045] The server then parses the received text information and uses a natural language processing (NLP) model (e.g., BERT, GPT-3) to understand its meaning and structure. The text information is converted into a visually understandable format, which is then transformed into visual content in the form of images or videos using a generative AI model (e.g., GAN, DALL-E).

[1046] The server then customizes the generated visual content based on the user's profile information (e.g., type of visual impairment, customization preferences, etc.), for example, by highlighting certain colors or adjusting font size. The customized visual content is temporarily stored in a database within the server.

[1047] Finally, the server transmits the customized visual content to the smart glasses via a communication network, and the smart glasses display the received visual content for the user to view in real time.

[1048] As a concrete example, consider a user with dyslexia capturing a product label in the detergent section of a brick-and-mortar store. The user uses the camera in their smart glasses to capture the label: "Detergent - 750 yen special offer!"

[1049] 1. The camera takes a picture of the label and uses OCR technology (e.g., Tesseract OCR) to extract the text information: "Detergent - Special Price: 750 yen!"

[1050] 2. This text information is sent from the smart glasses to the server.

[1051] 3. The server uses an NLP model (e.g., GPT-3) to convert the text information into a visually understandable format, and a generative AI model (e.g., GAN) to generate a customized image (e.g., highlighting the "On Sale!" part in red).

[1052] 4. The server sends the generated visual content to the smart glasses, and the user can view the received visual content in real time.

[1053] A specific example of a prompt is "Please convert the following text into a visually understandable format. Text: 'Detergent - Special Offer for 750 yen!'" Based on this prompt, the server converts the text information using an NLP model and then generates visual content that is easy to understand using a generative AI model.

[1054] This allows users with dyslexia to easily understand products and information in physical stores, allowing them to shop and use facilities comfortably.

[1055] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1056] Step 1:

[1057] The user wears smart glasses and uses the camera to capture product labels and information signs in a physical store. The input is image data captured by the smart glasses' camera. The output is saved as image data on the device. Specifically, the user points the camera at the information they want to see and presses a button to capture the image.

[1058] Step 2:

[1059] The device extracts text information from the captured image using OCR (Optical Character Recognition) technology (e.g., Tesseract OCR). The input is the captured image data. The output is the extracted text information. Specifically, the OCR software analyzes the image, recognizes the text, and extracts it as character data.

[1060] Step 3:

[1061] The extracted text information is sent from the device to the server via a communication network. The input is the text information extracted by the OCR. The output is the text information received by the server. Specifically, the device uploads the text information to the server via an Internet connection.

[1062] Step 4:

[1063] The server uses an NLP model (e.g., BERT, GPT-3) to parse the received text information. The input is the text information sent from the device. The output is the text information converted to make it easier to understand visually. Specifically, the NLP model analyzes the meaning and structure of the text and performs natural language processing.

[1064] Step 5:

[1065] The server uses a generative AI model (e.g., GAN, DALL-E) to generate images and videos based on the information obtained using the NLP model and converts it into a visually easy-to-understand format. The input is text information analyzed by the NLP model. The output is images or videos in a visually easy-to-understand format. Specifically, the generative AI model creates appropriate visual content.

[1066] Step 6:

[1067] The server customizes the generated visual content based on the user's profile information. The input is images or videos converted to be visually easy to understand, along with the user's profile information. The output is customized visual content. Specific actions include customization to suit the user's needs, such as enhancing colors or adjusting font sizes.

[1068] Step 7:

[1069] The server transmits the customized visual content to the user's smart glasses via a communication network. The input is the customized visual content. The output is the visual content received by the smart glasses. Specific operations include data transfer from the server to the smart glasses.

[1070] Step 8:

[1071] The terminal displays the received visual content, and the user can view it in real time. The input is the customized visual content sent from the server. The output is the information that the user visually views. In concrete terms, the visual content is displayed on the display of the smart glasses, allowing the user to view the information.

[1072] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1073] The system of the present invention supports users with dyslexia in visually understanding textual information in their daily lives and studies. This embodiment also adds a function to recognize the user's emotions and optimize visual content based on those emotions. This system involves a series of processes: the user's device acquires textual information, the server converts it into customized visual content, and the server further optimizes it based on the user's emotions and transmits it to the visual device.

[1074] System configuration and program handling

[1075] 1. Obtaining text information

[1076] A user uses a device such as a smartphone or AR glasses to capture text information to be read. For example, they take a picture of a sign or a page in a textbook with a camera. The device then extracts the text information from the captured image using OCR (optical character recognition) technology. The extracted text information is then sent to a server via a communication network.

[1077] 2. Analyzing textual information and converting it into visual content

[1078] The server receives the text information sent from the device and performs syntactic analysis. Syntactic analysis understands the meaning and structure of the text. The server then uses a natural language processing (NLP) model to convert the text information into a visually understandable format. This converted text information is then further converted into images or videos using generative AI (e.g., GAN: Generative Artificial Intelligence Network).

[1079] 3. Individual customization

[1080] The server customizes the generated visual content based on the user's profile information (type of visual impairment, customization preferences, etc.) For example, it may emphasize certain colors for color-blind users, or adjust font size and placement to make them easier to understand. The customized visual content is temporarily stored in a database on the server.

[1081] 4. Optimization by Emotion Engine

[1082] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's facial expression data or voice data to identify the emotion the user is currently feeling (e.g., joy, sadness, surprise, anger, etc.). Based on this emotion data, the server further adjusts the customization of the visual content. For example, if the user is feeling stressed, the server may change the color tone to one that is visually more relaxing.

[1083] 5. Transfer and display on a visual device

[1084] The server sends the customized visual content to the user's device or visual device (e.g., AR glasses or VR headset). This transfer occurs in real time, allowing the user to instantly view the information. The device then displays the received visual content, allowing the user to visually interpret the text information.

[1085] Specific examples

[1086] When reading signs in the city

[1087] The user points their smartphone camera at a sign and captures it. The device extracts text from the sign image and sends it to the server. The server analyzes the received text information and uses generative AI to convert the text into visual content. The server then customizes the content based on the user's profile information and uses an emotion engine to recognize and optimize the user's emotional state. The customized visual content is then sent to the device, allowing the user to visually understand information about the city.

[1088] When reading the contents of a textbook

[1089] A student user takes a photo of a textbook page with their smartphone. The device extracts text from the textbook image and sends it to the server. The server analyzes the received text information and converts it into a visually understandable format. The server then customizes the content based on the user's profile information, and an emotion engine analyzes the user's concentration level and emotional state to optimize the visual content. The user can view the textbook content in an easy-to-understand format through the adjusted visual content.

[1090] In this way, the system of the present invention not only makes it easier for users with dyslexia to visually understand textual information, but also provides visual content that is optimally adjusted according to the user's emotional state.

[1091] The processing flow will be explained below.

[1092] Step 1:

[1093] A user uses a device such as a smartphone or AR glasses to capture text information to be read. For example, the user takes a picture of a sign or a page in a textbook with a camera. The device then acquires the captured image data.

[1094] Step 2:

[1095] The device extracts text information from the acquired image data using OCR (Optical Character Recognition) technology, which converts the characters in the image into digital text. The device then sends the extracted text information to the server.

[1096] Step 3:

[1097] The server receives the text information sent from the terminal, analyzes the received text information, and performs syntax analysis. This analysis allows the meaning and structure of the text to be understood.

[1098] Step 4:

[1099] The server uses a natural language processing (NLP) model to convert the parsed text information into a visually understandable format, and then invokes a generative AI (e.g., GAN: Generative Artificial Intelligence Network) to convert the text into visual content such as images or videos, thereby generating the text information in a visual format.

[1100] Step 5:

[1101] The server references the user's profile information, which includes the user's type of visual impairment and customization preferences, and customizes the generated visual content based on this information. For example, it may enhance certain colors for users with color blindness, or change font size or layout.

[1102] Step 6:

[1103] The server recognizes the user's emotions using an emotion engine, which analyzes the user's facial expression data or voice data to identify the emotion the user is currently feeling (e.g., joy, sadness, surprise, anger, etc.).

[1104] Step 7:

[1105] The server further adjusts the visual content based on the user's emotional state, for example, changing the color tones to be more visually relaxing if the user is feeling stressed.

[1106] Step 8:

[1107] The server then sends the customized visual content to the user's device or viewing device (e.g., AR glasses or VR headset). This transfer occurs in real time, allowing the user to view the information instantly.

[1108] Step 9:

[1109] The device displays the received visual content. By looking at the visual content generated through the device or visual device, the user can visually understand the text information they have read. This makes it easier for people with disabilities, such as dyslexia, to understand the information.

[1110] Example 2

[1111] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1112] Conventional visual content generation systems focus on converting text information into something visually easy to understand, but do not fully consider the visual characteristics and emotional state of each user. This makes it difficult to provide optimal visual information, especially for users with dyslexia. Furthermore, they lack flexible customization to allow users to obtain information without feeling stressed or burdened.

[1113] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring text information, means for converting the text information into images or videos, means for customizing the images or videos based on the user's profile information, means for analyzing the user's emotional state and optimizing the visual content, and means for transmitting the customized and optimized images or videos to the user's visual device. This makes it possible to provide optimized visual content that takes into account the visual characteristics and emotional state of each user.

[1114] "Text information" is data of characters and sentences that the user should read.

[1115] The "means for acquiring" is a mechanism for acquiring text information using the user's device.

[1116] The "means for converting into images or videos" is a mechanism for converting acquired text information into images or videos that are visually easy to understand.

[1117] "User profile information" is data about an individual user's visual characteristics, preferences, and disabilities.

[1118] A "means for customizing" is a mechanism for optimizing and tailoring generated visual content to a particular user based on the user's profile information.

[1119] The "emotional state of the user" refers to the mood or state of mind of the user that can be understood through detected facial expressions and voice data.

[1120] The "analyzing means" is a mechanism for analyzing facial expression data and voice data to understand the emotional state of the user.

[1121] The "optimizing means" is a mechanism that further adjusts the visual content based on the user's emotional state to improve user comfort.

[1122] A "visual device" is a device used by a user to view visual content, including, for example, an augmented reality device or a virtual reality device.

[1123] A "transmitting means" is a mechanism that transfers customized and optimized visual content to a user's visual device.

[1124] A "system" is a collection of devices and software that include each of these means and operate in conjunction with one another.

[1125] The system of the present invention is designed to help users with dyslexia visually understand written information in their daily lives and studies. This system involves a series of processes: acquiring the user's text information, analyzing and converting it on the server, and further customizing and optimizing it based on the user's profile information and emotional state.

[1126] Specific examples of hardware and software used

[1127] Hardware:

[1128] Smartphone

[1129] Augmented Reality (AR) Devices

[1130] Virtual Reality (VR) Devices

[1131] software:

[1132] OCR (optical character recognition) technology

[1133] Parsing Engine

[1134] Natural Language Processing (NLP) Models

[1135] Generative AI (e.g., GAN: Generative Adversarial Networks)

[1136] Sentiment Analysis Engine

[1137] System operation example

[1138] When reading signs in the city

[1139] The user points their smartphone camera at a sign and captures it. The device uses OCR technology to extract text information, which is then sent to the server. The server then parses the received text and uses generative AI to convert it into visually understandable images or videos. The server then customizes the visual content based on the user's profile information and optimizes it using an emotion engine to recognize the user's emotional state. The customized and optimized visual content is then sent to the device, allowing the user to instantly visually understand the information on signs around town.

[1140] When reading the contents of a textbook

[1141] A student user takes a photo of a textbook page with their smartphone. The device uses OCR technology to extract text from the textbook image and sends it to the server. The server then parses the received text information and uses generative AI to convert it into a visually understandable format. The server then customizes the visual content based on the user's profile information and optimizes it using an emotion engine to analyze the user's concentration level and emotional state. The user can visually understand the content of the textbook through the tailored visual content.

[1142] Prompt Sentence Examples

[1143] "Use GANs to run a program that translates text information into visual content for dyslexic users, and optimize it for the user's emotional state. For example, if the user is stressed, adjust the color tones to be more relaxing."

[1144] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1145] Step 1: Obtaining text information

[1146] A user uses a device such as a smartphone or AR glasses to capture text information using the camera. For example, the user takes a picture of a sign in the city or a page in a textbook.

[1147] Input: Image data taken by the user

[1148] The device inputs the image data into OCR (optical character recognition) software to extract text information.

[1149] Output: Extracted text information

[1150] The terminal compresses the extracted text information, performs an error check, encrypts it, and transmits it to the server.

[1151] Step 2: Parsing and transforming text information

[1152] The server receives the text information sent from the terminal and stores it in a database.

[1153] Input: Text information sent from the device

[1154] The server uses natural language processing (NLP) models to analyze the text and understand its meaning and context, including morphological analysis and dependency analysis.

[1155] Output: Parsed text information

[1156] The server uses generative AI (e.g., GAN) based on the analyzed text information to convert it into visually easy-to-understand image or video content.

[1157] Output: The generated visual content (images or videos)

[1158] Step 3: Individual customization

[1159] The server refers to the user's profile information (visual characteristics, preferences, etc.) stored in a database.

[1160] Input: Generated visual content, user profile information

[1161] The server customizes the visual content based on the user's profile information, for example by enhancing certain colors for color-blind users or adjusting font size and placement.

[1162] Output: Customized visual content

[1163] The server temporarily stores the customized visual content in a database.

[1164] Step 4: Optimizing with an Emotional Engine

[1165] The server inputs facial expression data and voice data acquired from the terminal into an emotion engine and analyzes the user's emotional state.

[1166] Input: facial expression data, voice data

[1167] The server uses an emotion engine to identify the user's emotion (joy, sadness, surprise, anger, etc.).

[1168] Output: Emotion analysis results

[1169] Based on the results of the emotion analysis, the server adjusts the color tone and layout of the visual content, optimizing it so that users can view information more comfortably.

[1170] Output: Optimized visual content

[1171] Step 5: Transfer and display to a visual device

[1172] The server transmits optimized visual content in real time to the user's device or visual device (such as AR glasses or VR headset).

[1173] Input: Optimized visual content

[1174] The terminal decodes the received visual content and displays it on its display.

[1175] Output: The displayed visual content

[1176] Users view customized and optimized visual content through their terminals and visual devices and visually understand textual information.

[1177] (Application example 2)

[1178] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1179] Traditionally, users with dyslexia have had difficulty visually understanding textual information. This can also hinder their ability to understand textual information during everyday activities such as online shopping and information searches, resulting in stress and misunderstandings. Furthermore, visual content provided without considering the user's emotional state can make comprehension even more difficult. The present invention aims to solve these problems and enable users with dyslexia to easily understand textual information and live their daily lives comfortably.

[1180] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring text information, means for analyzing the acquired text information and converting it into a form that is visually easy to understand, means for recognizing a user's emotion and optimizing visual content based on the emotion, means for customizing the visual content based on profile information, and means for transmitting the customized and optimized visual content to the user's terminal or visual device. This makes it easier for users with dyslexia to visually understand text information, and further enables the provision of optimal content according to their emotional state.

[1181] "Text information" is data of characters or sentences that are visually displayed.

[1182] "Means for acquiring text information" refers to a method of capturing visual characters or sentences as electronic data using a device such as a camera or scanner.

[1183] "Analysis" refers to processing the acquired text information to understand its meaning and structure.

[1184] "Means of converting into a visually easy-to-understand form" refers to methods of converting text information into a visually easy-to-understand form such as an image or video.

[1185] "Means for recognizing a user's emotions and optimizing visual content based on those emotions" refers to a method for analyzing a user's current emotional state from facial expressions and voice, and adjusting visual content to appropriately reflect those emotions.

[1186] "Profile information" refers to information such as the type of visual impairment and customization preferences of each individual user.

[1187] "Means for customizing visual content" refers to methods for adjusting color settings, font size, placement, etc. to suit each user based on profile information.

[1188] "Visual devices" are devices for visually displaying digital information, such as AR glasses and VR headsets.

[1189] A "server" is a centralized computer system for processing and managing data.

[1190] A "user's terminal" is a device that a user can operate at hand, such as a smartphone or tablet.

[1191] A "natural language processing model" is a learning model for analyzing and generating text data.

[1192] An "emotion engine" is software for analyzing a user's emotional state.

[1193] A "generative AI model" is an artificial intelligence model that uses neural networks to generate new data.

[1194] This invention provides a system that helps users with dyslexia visually understand text information in daily life and online shopping. In this embodiment, this system is realized through the following steps.

[1195] The server analyzes the text information sent by the user and converts it into a visually understandable format. First, the user captures the text information they want to obtain using a device such as a smartphone or tablet. From this captured image, the text information is extracted using optical character recognition (OCR) technology. The extracted text information is then sent to the server.

[1196] The server performs syntactic analysis on the received text to understand its meaning and structure, then uses a natural language processing (NLP) model to convert the text into a visually understandable format, which is then converted into an image or video using a generative AI model (e.g., a generative neural network).

[1197] The server then customizes the visual content based on the profile information, which may include, for example, the type of visual impairment and customization preferences, such as color enhancement or font size adjustment.

[1198] Additionally, the server uses an emotion engine to recognize the user's emotional state. It analyzes the user's facial expressions and voice data to identify their current emotion (happiness, sadness, stress, etc.). Based on the recognized emotion, the visual content is optimized. For example, if the user is feeling stressed, the visual color tone will be changed to a more relaxing one.

[1199] The customized and optimized visual content is then sent from the server to the user's terminal or visual device, such as an augmented reality (AR) glass or a virtual reality (VR) headset, which then displays the received visual content in real time, allowing the user to instantly visualize the information.

[1200] For example, when a user with dyslexia tries to understand product information on an online shopping website, the process is as follows: The user uses their smartphone camera to capture the product description. The device uses OCR technology to extract text from the image and sends it to the server. The server then analyzes the text and converts it into a visually understandable format. This converted information is customized based on the user's profile information and further optimized for the user's emotional state. The customized visual content is then sent to the user's device, where it can be transparently understood.

[1201] Examples of prompts:

[1202] "Please take a picture of this product description to make it easier to understand visually."

[1203] In this way, the system of the present invention can provide textual information in a visually easy-to-understand format when a user with dyslexia is shopping online, and can also provide visual content optimized according to the user's emotional state.

[1204] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1205] Step 1:

[1206] A user uses a smartphone or tablet to capture text information to read. Specifically, they open a camera app and take a picture of a product description or review. The input is the image taken with the smartphone camera, and the output is the captured image data.

[1207] Step 2:

[1208] The device extracts text information from the acquired image data using OCR technology. In this process, the image data is input, and an OCR engine (e.g., Tesseract) is used to recognize characters, resulting in text data as the output.

[1209] Step 3:

[1210] Text data is sent from a terminal to a server. The input is text data that is transferred to the server via a communication network. The output is the text data that arrives at the server.

[1211] Step 4:

[1212] The server analyzes the received text data and converts it into a form that is easy to understand visually. Specifically, it performs semantic analysis using a natural language processing model (e.g., BERT or GPT) and generates images and videos using a generative AI model (e.g., GAN). The input is text data and the output is visual content (images and videos).

[1213] Step 5:

[1214] The server customizes the visual content based on the user's profile information, for example by adjusting colors to accommodate color blindness or changing text size. The inputs to this process are the user's profile information and the visual content, and the output is the customized visual content.

[1215] Step 6:

[1216] The server uses an emotion engine to recognize the user's emotional state. Specifically, it analyzes facial expression images and voice data provided by the user to identify emotions. The input of this process is emotion data (facial expression images and voice data), and the output is the recognized emotional state.

[1217] Step 7:

[1218] The server further optimizes the visual content based on the recognized emotional state, for example, changing the color tones to a more relaxing color if the user is stressed. The inputs to this process are the emotional state and the customized visual content, and the output is the optimized visual content.

[1219] Step 8:

[1220] The optimized visual content is sent from the server to the user's device or visual device. The input is the optimized visual content, which is sent to the user's device or AR glasses via a communication network. The output is the visual content displayed on the user's device or visual device.

[1221] Step 9:

[1222] The user checks the visual content displayed on the terminal or visual device and visually understands the required information. The input is the optimized content displayed on the visual device, and the output is the user's understanding. This step allows the user to visually understand the information comfortably.

[1223] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1224] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1225] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1226] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1227] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1228] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1229] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1230] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1231] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1232] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1233] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1234] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1235] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1236] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1237] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1238] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1239] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1240] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1241] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1242] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1243] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1244] The following is further disclosed regarding the above embodiment.

[1245] (Claim 1)

[1246] a means for obtaining text information;

[1247] A means for converting the acquired text information into an image or video;

[1248] means for customizing the image or video based on user profile information;

[1249] means for transmitting the customized image or video to the user's visual device;

[1250] A system including:

[1251] (Claim 2)

[1252] 10. The system of claim 1, wherein the visual device is an augmented reality (AR) device or a virtual reality (VR) device.

[1253] (Claim 3)

[1254] 10. The system of claim 1, wherein the means for obtaining text information uses optical character recognition (OCR) technology.

[1255] "Example 1"

[1256] (Claim 1)

[1257] a means for obtaining text information;

[1258] means for transmitting the acquired text information to a server via a communication network;

[1259] a means by which the server parses and converts the textual information into visual content;

[1260] A means of converting it into a visually understandable form using a generative AI model, and

[1261] means for customizing said visual content based on user profile information;

[1262] means for transmitting customized visual content to a user's visual device;

[1263] A system including:

[1264] (Claim 2)

[1265] 10. The system of claim 1, wherein the visual device is an augmented reality (AR) device or a virtual reality (VR) device.

[1266] (Claim 3)

[1267] 10. The system of claim 1, wherein the means for obtaining text information uses optical character recognition (OCR) technology.

[1268] "Application Example 1"

[1269] (Claim 1)

[1270] a means for obtaining text information;

[1271] A means for converting the acquired text information into an image or video;

[1272] means for customizing the image or video based on user profile information;

[1273] means for transmitting the customized image or video to the user's visual device;

[1274] A method to capture product labels and guide signs in physical stores using smart glasses and convert the text information into a visually easy-to-understand format.

[1275] means for displaying the generated visual content on the smart glasses so that the user can view the visual information in real time;

[1276] A system including:

[1277] (Claim 2)

[1278] 10. The system of claim 1, wherein the visual device is an augmented reality (AR) device or a virtual reality (VR) device.

[1279] (Claim 3)

[1280] 10. The system of claim 1, wherein the means for obtaining text information uses optical character recognition (OCR) technology.

[1281] "Example 2: Combining Emotion Engines"

[1282] (Claim 1)

[1283] a means for obtaining text information;

[1284] A means for converting the acquired text information into an image or video;

[1285] means for customizing the image or video based on user profile information;

[1286] means for analyzing a user's emotional state and optimizing visual content;

[1287] means for transmitting the customized and optimized image or video to the user's visual device;

[1288] A system including:

[1289] (Claim 2)

[1290] 10. The system of claim 1, wherein the visual device is an augmented reality (AR) device or a virtual reality (VR) device.

[1291] (Claim 3)

[1292] 10. The system of claim 1, wherein the means for obtaining text information uses optical character recognition (OCR) technology.

[1293] "Application example 2 when combining emotion engines"

[1294] (Claim 1)

[1295] a means for obtaining text information;

[1296] A means to analyze the acquired text information and convert it into a visually understandable format,

[1297] means for recognizing a user's emotion and optimizing visual content based on the emotion;

[1298] means for customizing visual content based on profile information;

[1299] means for transmitting customized and optimized visual content to a user's terminal or visual device;

[1300] A system including:

[1301] (Claim 2)

[1302] 10. The system of claim 1, wherein the system is an augmented reality device or a virtual reality device.

[1303] (Claim 3)

[1304] 10. The system of claim 1, which uses optical character recognition technology. [Explanation of symbols]

[1305] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for obtaining text information; A means for converting the acquired text information into an image or video; means for customizing the image or video based on user profile information; means for transmitting the customized image or video to the user's visual device; A system including:

2. The system of claim 1 , wherein the visual device is an augmented reality (AR) device or a virtual reality (VR) device.

3. 10. The system of claim 1, wherein the means for obtaining text information uses optical character recognition (OCR) technology.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A