System

The visual assistance system addresses the challenge of visual field loss by using a camera, AI server, and retinal projection to provide real-time compensation for visual defects, improving safety and convenience for visually impaired individuals.

JP2026016227APending Publication Date: 2026-02-03SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024117317
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Patients with visual field narrowing due to glaucoma or other conditions often miss important visual information, posing safety risks in daily activities like walking or driving, as conventional visual assistance devices lack real-time compensation for specific visual information.

Method used

A visual assistance system comprising a camera, AI analysis server, and retinal projection device that captures, analyzes, and projects corrected video data onto the user's retina, recognizing objects and text, and reconstructing images to compensate for visual field defects.

Benefits of technology

The system effectively complements visual information in real-time, enabling visually impaired users to safely navigate their environment by projecting important information directly onto the retina, enhancing safety and convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016227000001_ABST
    Figure 2026016227000001_ABST
Patent Text Reader

Abstract

To provide a vision support system for effectively supporting a visual field defect such as glaucoma and improving the quality of daily life.SOLUTION: A vision aid system comprising: means for capturing an image by a camera; a server having a AI of analyzing the captured image; and a projector for correcting the analyzed image and projecting the corrected image onto retinas of a user. Specifically, a means for recognizing an object in an image, a means for extracting character information, and a means for reconstructing an image by correcting a visual field defect portion are provided, and acquisition of visual information is supported by generating and projecting image data corrected in real time on the retina of the user. In addition, it is possible to improve security and convenience by rearranging important information in a video in a form that is easily recognized by a user through AI analysis and appropriately displaying text information or object information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Patients who experience visual field narrowing due to glaucoma or other conditions, resulting in loss of certain visual fields, are likely to miss important visual information in their daily lives, posing a particular threat to safety when walking or driving. Conventional visual assistance devices have technical limitations in compensating for specific visual information, making it particularly difficult to provide visual information in real time. The present invention aims to provide a visual assistance system that effectively compensates for visual field loss due to glaucoma and other conditions, thereby improving the quality of daily life. [Means for solving the problem]

[0005] This invention provides a visual assistance system that includes a camera for capturing video, a server equipped with AI for analyzing the captured video, and a projection device for correcting the analyzed video data and projecting it onto the user's retina. Specifically, the system includes a means for recognizing objects in the video, a means for extracting text information, and a means for correcting visual field defects and reconstructing the video. The system generates and projects corrected video data onto the user's retina in real time to support the acquisition of visual information. Furthermore, AI analysis can rearrange important information in the video in a way that is easy for the user to recognize, and appropriately display text and object information, thereby improving safety and convenience.

[0006] A "camera" is a device that captures video and outputs it as a digital signal.

[0007] "Video" is a collection of consecutive frames of visual information captured by a camera.

[0008] "AI" is an abbreviation for artificial intelligence and refers to algorithms and models for analyzing data and performing object and character recognition.

[0009] A "server" is a computer system that processes data and provides services to other devices over a network.

[0010] A "projection device" is a device for displaying analyzed image data onto the user's retina.

[0011] The term "visual aid system" refers to the entire system that complements and enhances visual information, and its purpose is to compensate for visual field defects.

[0012] "Object recognition" is the process of identifying specific objects contained in a video and extracting information related to them.

[0013] "Text information" refers to the text data contained in the video, and includes the process of analyzing and reading it.

[0014] "Visual field defect" refers to the phenomenon in which part of the visual field is lost, resulting in an area that cannot be seen.

[0015] "Reconstruction" is the process of generating new images based on analyzed data to complement visual information.

[0016] "User" refers to a person using a visual aid system, particularly a patient with visual field defects. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] An embodiment of this invention provides a system that complements visual information for users with impaired vision due to glaucoma, etc. This system is centered around a glasses-type device worn by the user, and is composed of a camera, an AI analysis server, and a retinal projection device.

[0039] System configuration

[0040] 1. Camera:

[0041] It is installed in glasses worn by the user and captures images in front of the vehicle in real time, which are then saved as image frame data.

[0042] 2. AI analysis server:

[0043] The system receives video data sent from the camera and analyzes it using an AI model. Specifically, it recognizes objects and characters in the video and identifies the information.

[0044] To compensate for the visual field defects, the analysis results are used to generate corrected video data, which includes scaling the video and adding text information.

[0045] 3. Retinal projection device:

[0046] This device projects corrected image data onto the user's retina in real time, allowing the user to compensate for visual field defects and easily obtain visual information in everyday life.

[0047] Program processing (natural language explanation)

[0048] Video capture and transmission

[0049] Device (camera): The camera in the glasses worn by the user captures the image in front of the user in real time. This image is saved as data frame by frame.

[0050] Terminal: Sends captured video data to an AI analysis server via the internet.

[0051] Video data analysis

[0052] Server: The AI ​​analytics server processes the received video data in real time, using AI models to recognize objects in the video, such as cars, park benches, and traffic lights, and assigns bounding boxes to each.

[0053] Server: Using OCR technology, character information in the video is extracted, obtaining text data such as "Watch out for pedestrians" or "The traffic light is red."

[0054] Video data correction

[0055] Server: Based on the analysis results, an image is generated to compensate for the visual field defect. Important information is reconstructed in a form that is easy for the user to recognize, using techniques such as image enlargement.

[0056] Server: Object information and extracted text information are overlaid on the video in a simple format.

[0057] Video data transmission and projection

[0058] Server: Sends the corrected video data to the user terminal.

[0059] Terminal: The transmitted corrected image data is projected onto the user's retina in real time. This process complements the missing information in the user's field of vision.

[0060] Specific examples

[0061] 1. If the user is waiting at a traffic light:

[0062] Terminal (camera): Captures the traffic light in front of the user.

[0063] Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, extracts the traffic light's text information (e.g., "pedestrian signal").

[0064] Server: Reconstructs the traffic light color and text information as a corrected image.

[0065] Device: Projects traffic light colors and text information onto the user's retina to complement visual information.

[0066] 2. If the user is walking down the street:

[0067] Terminal (camera): Captures images of the road and intersections ahead.

[0068] Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[0069] Server: Emphasizes information about pedestrians and vehicles that require attention and generates corrected images.

[0070] Terminal: This information is projected onto the user's retina to complement the visual information.

[0071] In this way, the system effectively supports the vision of glaucoma patients, helping them live safe and comfortable lives.

[0072] The processing flow will be explained below.

[0073] Step 1:

[0074] Device (camera): A camera mounted on glasses worn by the user captures images of the surroundings in real time, and these images are stored as digital data frame by frame.

[0075] Step 2:

[0076] Terminal: Captured video data is sent to the AI ​​analysis server via the internet using a low-latency, highly reliable communication protocol.

[0077] Step 3:

[0078] Server: Stores the received video data in a buffer for analysis.

[0079] Step 4:

[0080] Server: The AI ​​model analyzes the video data stored in the buffer and recognizes objects in the image, such as cars, pedestrians, and traffic lights, and assigns labels and bounding boxes to each.

[0081] Step 5:

[0082] Server: Analyzes text information contained in video using OCR technology and saves it as metadata. For example, extracts text from traffic lights and signs.

[0083] Step 6:

[0084] Server: Based on the analysis results, an image is generated that corrects for the visual field defect. Here, processing such as enlarging parts of the image is performed to make important information easier for the user to see.

[0085] Step 7:

[0086] Server: Object information and extracted text information are overlaid on the original video. For example, a large "STOP" message is displayed over a "red light."

[0087] Step 8:

[0088] Server: Converts the corrected video data into a format suitable for head-up displays and sends it to the user's device.

[0089] Step 9:

[0090] Terminal: The received corrected image data is projected onto the user's retina in real time, allowing the user to fill in information about the missing area of ​​their visual field.

[0091] Step 10:

[0092] User: Views the projected image and recognizes the surrounding situation and objects, enabling safe walking and driving.

[0093] Through this series of steps, the visual aid system effectively compensates for the user's visual field defects and provides the visual information necessary for daily life.

[0094] Example 1

[0095] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0096] Visually impaired users face challenges in moving safely and accurately grasping information about their surroundings in daily life. In particular, if they have lost vision due to diseases such as glaucoma, they are more likely to miss important visual information in the lost area, which could lead to serious accidents or danger.

[0097] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0098] In this invention, the server includes means for recognizing objects in the video, means for extracting text information, means for correcting the visual field defect and reconstructing the video, and means for enlarging the corrected video and adding text information, thereby enabling visually impaired users to supplement important information present in the visual field defect in real time and recognize their surroundings safely and reliably.

[0099] A "camera" is a photographic device for capturing images.

[0100] "Video data" refers to visual information captured by a camera and recorded in digital format.

[0101] "Artificial intelligence" is a technology in which computer programs mimic human intelligence to solve problems, and in this case it is used to analyze video data.

[0102] A "server" is a computer device that receives data over a network and analyzes and processes it.

[0103] A "projection device" is a device for projecting the analyzed and corrected image data onto the user's retina.

[0104] A "network" is an infrastructure for data communication, and is used here to transmit data between the camera and the server.

[0105] "Object recognition" is a technology that detects specific objects in video data and identifies their location and type.

[0106] "Text information" refers to characters and text data contained in video data.

[0107] "OCR" stands for Optical Character Recognition, a technology that extracts text information from video data.

[0108] The "visual field defect area" refers to an area that is missing from the user's field of vision due to a visual impairment.

[0109] "Image correction" is a correction process carried out to compensate for missing parts of the field of view based on analyzed image data.

[0110] "Real-time" refers to processing and display occurring immediately, without delay.

[0111] "Image enlargement" is a technique for displaying a specific image portion in a larger, more easily visible manner.

[0112] "Adding text information" is a technique for overlaying analyzed text information onto video.

[0113] This invention provides a system that complements visual information for users with visual impairments due to glaucoma, etc. The system consists of a camera, an AI analysis server, a network, and a projection device.

[0114] 1. Camera: The camera is mounted on a glasses-type device worn by the user and captures the image in front of the user in real time. The captured image is recorded as digital data frame by frame.

[0115] 2. AI analysis server: The video data captured by the camera is sent to the AI ​​analysis server via the network. The server performs the following processing steps:

[0116] Object Recognition: Using AI models to identify objects in a video (cars, traffic lights, pedestrians, etc.) and assign them bounding boxes, for example, identifying a traffic light as red.

[0117] Text information extraction: Extract text information (e.g., "pedestrian signal") from the video using OCR technology.

[0118] Image correction: To generate a corrected image, areas of vision defects are filled in and important information is emphasized, such as overlaying traffic light colors or important messages on the image.

[0119] 3. Network: Wi-Fi or 4G / 5G networks are used for data communication, sending and receiving data between the camera and the server.

[0120] 4. Projection device: Projects corrected image data onto the user's retina in real time, allowing the user to instantly see important information present in the area of ​​visual field loss.

[0121] Specific examples

[0122] 1. If the user is waiting at a traffic light:

[0123] Terminal (camera): Captures the traffic light in front of the user.

[0124] Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, extracts the traffic light's text information (e.g., "pedestrian signal").

[0125] Server: Reconstructs the traffic light color and text information as a corrected image.

[0126] Device: Projects traffic light colors and text information onto the user's retina to complement visual information.

[0127] 2. If the user is walking down the street:

[0128] Terminal (camera): Captures images of the road and intersections ahead.

[0129] Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[0130] Server: Emphasizes information about pedestrians and vehicles that require attention and generates corrected images.

[0131] Terminal: This information is projected onto the user's retina to complement the visual information.

[0132] Prompt Sentence Examples

[0133] Traffic light discrimination prompt: The prompt for extracting color and text information from traffic light images, generating corrected images for retinal projection is as follows: "Please distinguish between red, yellow, and green signals from traffic light images, extract the traffic light text information, and generate corrected images. Then, send the generated images to the retinal projection device."

[0134] Object Recognition Prompt: The prompt for recognizing cars, pedestrians, and traffic lights at intersections from the forward video while the user is walking and generating corrected video for safety information is as follows: "Please recognize cars, pedestrians, and traffic lights from the forward video while the user is walking, and reconstruct the corrected video for safety information. Then, send the generated video to the retinal projection device."

[0135] This system will enable glaucoma patients and other visually impaired users to effectively supplement their visual information in their daily lives, enabling them to live safe and comfortable lives.

[0136] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0137] Step 1:

[0138] The user turns on the eyeglasses-type device.

[0139] Input: Power-on operation

[0140] Specific operation: The user presses a button to turn on the device, which starts up the device's system and initializes the camera and retinal projection device.

[0141] Output: Device booted, camera and retinal projection device ready

[0142] Step 2:

[0143] The device (camera) captures the image in front of the user.

[0144] Input: Device boot complete, user visible

[0145] How it works: The camera is installed inside the device and captures real-time images of what is in front of the user's field of vision.

[0146] Output: Captured video data (frame by frame)

[0147] Step 3:

[0148] The device (camera) sends the captured video data to the server.

[0149] Input: Captured video data (frame by frame)

[0150] Specific operation: The captured video data is sent to an AI analysis server via Wi-Fi or 4G / 5G networks.

[0151] Output: Video data received by the server

[0152] Step 4:

[0153] The server inputs the received video data into the AI ​​model and begins analysis.

[0154] Input: Video data received by the server

[0155] How it works: The server inputs the video data into the AI ​​model and performs object recognition and OCR processing, such as identifying traffic lights, pedestrians, and text information and enclosing them in bounding boxes.

[0156] Output: Analyzed video data (object recognition information, text information)

[0157] Step 5:

[0158] The server generates an image that corrects the visual field defect based on the analysis results.

[0159] Input: Analyzed video data (object recognition information, text information)

[0160] How it works: The server uses the analysis results to generate an image that complements the visual field defect, enlarging important information and adding warning text such as "Watch your direction."

[0161] Output: Corrected video data

[0162] Step 6:

[0163] The server transmits the corrected video data to the user terminal.

[0164] Input: Corrected video data

[0165] Specific operation: The corrected video data is immediately sent to the user terminal and adjusted to minimize delay.

[0166] Output: Corrected video data received by the user device

[0167] Step 7:

[0168] The terminal (retinal projection device) projects the corrected image data onto the user's retina.

[0169] Input: Corrected video data received by the user terminal

[0170] How it works: The retinal projection device accurately projects the corrected image data into the user's field of vision, for example, displaying the red light of a traffic light and the text "Please stop" in the center of the field of vision.

[0171] Output: Corrected image projected onto the user's retina

[0172] In this way, each step of the system is performed sequentially, allowing visually impaired users to complete important visual information in real time.

[0173] (Application example 1)

[0174] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0175] This invention relates to a visual aid system that enables visually impaired users to appropriately acquire visual information and act safely in daily life and in specific situations. In particular, the objective is to provide a system that analyzes visual information in real time while driving a car and complements and emphasizes important information, thereby enabling visually impaired people to drive safely. Current visual aid systems lack analytical accuracy, real-time performance, and retinal projection technology, making improving driving safety a challenge.

[0176] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0177] In this invention, the server includes a means for capturing video using a camera, a computer equipped with artificial intelligence for analyzing the captured video, a projection device for correcting the analyzed video data and projecting it onto the user's retina, a means for transmitting the corrected video data from the computer to a terminal in real time, and a means for displaying the corrected video data on the terminal. This allows visually impaired users to accurately grasp visual information while driving, significantly improving safety. Furthermore, real-time analysis and visual enhancement allow users to immediately recognize important information such as obstacles and traffic signals, shortening the user's reaction time and making it easier to avoid danger.

[0178] A "camera" is a device for capturing images and is a piece of equipment used to record visual information.

[0179] "Artificial intelligence for video analysis" is a computer system that processes captured video data and performs object and character recognition.

[0180] A "computer" is a device that processes and analyzes data, and in particular, a computer equipped with artificial intelligence (AI)-based image analysis technology.

[0181] A "projection device" is a device that has the technology to project analyzed image data onto the user's vision, and in particular refers to equipment that projects images directly onto the retina.

[0182] "Real-time transmission means" refers to the technology and protocols for immediately transmitting processed data to the terminal.

[0183] A "terminal" is a device for displaying received video data, and includes smart glasses, head-mounted displays, and the like.

[0184] "Means for recognizing objects" refers to technology for detecting objects and people contained in video and identifying their location and type.

[0185] "Means for extracting text information" refers to OCR (optical character recognition) technology for analyzing and reading text data contained in video.

[0186] "Means for reconstructing images" refers to technology for correcting visual field defects and regenerating image data in a form that is easy for users to understand.

[0187] A "visual display device" is a device for visually presenting analyzed and corrected image data to a user, including smart glasses and retinal projection devices.

[0188] As an embodiment of the present invention, a system is provided for supplementing visual information for users with visual impairments such as glaucoma while driving a car. This system includes a camera, a computer (AI analysis server), a projection device, and a terminal. The role of each component and the overall processing flow are described in detail below.

[0189] System configuration

[0190] 1. Camera

[0191] The camera captures real-time images of the driving environment, which are then stored as data frame by frame and sent to a computer. The camera is used in smart glasses or head-mounted displays.

[0192] 2. AI analysis server

[0193] The computer (AI analysis server) receives the video data sent from the camera and analyzes it using an artificial intelligence model. Specifically, it recognizes objects in the video (cars, pedestrians, traffic lights, etc.) and extracts text information (intersection signs, road information, etc.). It then corrects for any visual defects and reconstructs the video in a way that emphasizes important information.

[0194] 3. Projection device

[0195] The corrected image data is projected directly onto the user's retina through a projection device, such as a device inside a pair of smart glasses, to complement the user's visual information in real time.

[0196] 4. Terminal

[0197] The corrected video data is sent in real time from the AI ​​analysis server to the device that displays the received video data, such as smart glasses or a head-mounted display.

[0198] Operational Overview

[0199] The specific steps involved in the operation of this system are described below.

[0200] 1. Video capture and transmission

[0201] The camera captures the image in front of the user and saves it as video frame data, which is then sent to an AI analysis server via the internet.

[0202] Example prompt sentence:

[0203] "Capture the video frame data to send to the AI ​​server and send it over the Internet. The video frame format is JPEG, and the server URL is 'http: / / ai-server / parse_frame'."

[0204] 2. Analysis and correction of video data

[0205] The server processes the received video data in real time. Using AI analysis, it recognizes objects in the video, identifying things like cars, park benches, and traffic lights. It also uses optical character recognition (OCR) technology to extract text information from the video. Based on the analysis results, it generates an image to complement any visual field defects, visually emphasizing important information.

[0206] Example prompt sentence:

[0207] "Please recognize objects and text contained in the received video data and return that information as a result. We use the YOLO model for object recognition and OCR technology for character recognition."

[0208] 3. Projecting the corrected image

[0209] The corrected image data sent from the server is received by the device and projected onto the user's retina in real time, allowing the user to supplement their visual information and accurately recognize important information while driving.

[0210] Example prompt sentence:

[0211] "Send the corrected image data to the retinal projection device so that the user can accurately perceive the visual information."

[0212] Specific examples

[0213] 1. Waiting at a traffic light

[0214] When the user approaches an intersection, the camera captures the traffic lights and intersection signs, and the AI ​​analysis server recognizes the color of the traffic lights (red, yellow, green) and provides that information to the user.

[0215] Example prompt sentence:

[0216] "Analyze the color and text information of traffic lights, and if the light is yellow, display 'Proceed with caution'."

[0217] 2. Walking on the street

[0218] As a user crosses the road, the camera captures the movements of pedestrians and vehicles. The AI ​​analytics server analyzes this information, highlights areas that require attention, and provides the user with enhanced video.

[0219] Example prompt sentence:

[0220] "Analyze the location of pedestrians and other vehicles as you approach an intersection and highlight any potentially dangerous situations."

[0221] This system will enable visually impaired people to obtain the visual information they need safely while driving a car, significantly improving driving safety.

[0222] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0223] Step 1:

[0224] A camera captures an image in front of the user.

[0225] Input: The user's current field of view.

[0226] Output: Video frame data.

[0227] How it works: The camera built into the smart glasses or head-mounted display captures video in real time and generates video frame data.

[0228] Step 2:

[0229] The device sends the captured video data to an AI analysis server.

[0230] Input: Video frame data.

[0231] Output: Video data sent to the server.

[0232] Specific operation: The device sends video data via the Internet, and the prompt to be entered is "Capture the video frame data to be sent to the AI ​​server and send it via the Internet. The video frame format is JPEG, and the server URL is 'http: / / ai-server / parse_frame'."

[0233] Step 3:

[0234] The server analyzes the received video data.

[0235] Input: Transmitted video data.

[0236] Output: Analysis results (object and character recognition data).

[0237] Specific operation: The server uses a generative AI model such as the YOLO model to recognize objects and text contained in the video data and identify information such as traffic lights, pedestrians, and signs. An example of a prompt is "Please recognize the objects and text contained in the received video data and return that information as a result. The YOLO model is used for object recognition, and OCR technology is used for character recognition."

[0238] Step 4:

[0239] The server corrects the video based on the analysis results.

[0240] Input: Analysis results.

[0241] Output: Corrected video data.

[0242] How it works: Based on the analysis results, the server compensates for the visual field defects and reconstructs the image to highlight important information. This correction makes the color of traffic lights and the content of signs appear larger, for example.

[0243] Step 5:

[0244] The server transmits the corrected video data to the terminal in real time.

[0245] Input: Corrected video data.

[0246] Output: Corrected video data sent to the device.

[0247] Specific operation: The server sends the corrected image data to the terminal in real time, and an example of a prompt is used: "Please send the corrected image data to the retinal projection device so that the user can accurately perceive the visual information."

[0248] Step 6:

[0249] The terminal displays the corrected image data on a display device (retinal projection device).

[0250] Input: Corrected video data.

[0251] Output: The corrected image displayed in the user's field of view.

[0252] How it works: The device sends the corrected image data to the retinal projection device, which projects it into the user's field of vision. For example, traffic light colors and sign information are projected directly onto the user's retina, allowing the user to perceive the necessary information, even if they have visual impairments.

[0253] By following the steps above, a visually impaired user can accurately obtain visual information while driving a car, enabling safe driving.

[0254] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0255] An embodiment of this invention provides a system that complements visual information for users with impaired vision due to glaucoma, etc., and also provides optimal auxiliary information according to the user's emotional state. This system is centered around a glasses-type device worn by the user, and is composed of a camera, an AI analysis server, an emotion engine, and a retinal projection device.

[0256] System configuration

[0257] 1. Camera:

[0258] The device is installed in glasses worn by the user and captures images of the surrounding area in real time, which are then stored as digital data frame by frame.

[0259] 2. AI analysis server:

[0260] The system receives video data sent from the camera and analyzes it using an AI model. Specifically, it recognizes objects and characters in the video and identifies the information.

[0261] To compensate for the visual field defects, the analysis results are used to generate corrected video data, which includes scaling the video and adding text information.

[0262] 3. Emotion Engine:

[0263] Recognize the user's emotional state in real time. The emotion engine can use technologies such as facial expression recognition and voice analysis.

[0264] Adjusting the display of visual information depending on emotional state (e.g., surprise, anxiety, joy, etc.).

[0265] 4. Retinal projection device:

[0266] This device projects corrected image data onto the user's retina in real time, allowing the user to compensate for visual field defects and easily obtain visual information in everyday life.

[0267] Program processing (natural language explanation)

[0268] Video capture and transmission

[0269] Device (camera): The camera in the glasses worn by the user captures the image in front of the user in real time, and this image is saved as digital data frame by frame.

[0270] Terminal: Sends captured video data to an AI analysis server via the internet.

[0271] Video data analysis

[0272] Server: The AI ​​analytics server processes the received video data in real time, using AI models to recognize objects in the video, such as cars, park benches, and traffic lights, and assigns labels and bounding boxes to each.

[0273] Server: Using OCR technology, the text information in the video is analyzed and text data such as "Watch out for pedestrians" or "The traffic light is red" is obtained.

[0274] Video data correction

[0275] Server: Based on the analysis results, an image is generated that corrects for the visual field defect. Here, processing such as enlarging parts of the image is performed to make important information easier for the user to see.

[0276] Server: Based on the object information and extracted text information, the server reconstructs the image with appropriate correction processing.

[0277] Emotion engine processing

[0278] Server: Analyzes the user's emotional state in real time using an emotion engine, for example, by using technology to detect emotions from facial expressions and voice.

[0279] Server: Based on the user's emotional state detected by the emotion engine, the server corrects the video data and optimizes the display content. For example, if the user is feeling anxious, it can highlight warning information.

[0280] Video data transmission and projection

[0281] Server: Transmits optimized video data based on correction and emotion information to the terminal.

[0282] Terminal: The received corrected image data is projected onto the user's retina in real time, allowing the user to compensate for the visual field defect and obtain appropriate visual information corresponding to their emotions.

[0283] Specific examples

[0284] 1. If the user is waiting at a traffic light:

[0285] Terminal (camera): Captures the traffic light in front of the user.

[0286] Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, analyzes the text displayed on the traffic lights.

[0287] Server: Generates corrected images based on the color and text information of traffic lights.

[0288] Server: If the emotion engine recognizes the user's state of anxiety, it highlights the traffic light information.

[0289] Device: Projects corrected and highlighted traffic light information onto the user's retina, allowing the user to see the traffic light with confidence.

[0290] 2. If the user is walking down the street:

[0291] Terminal (camera): Captures images of the road and intersections ahead.

[0292] Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[0293] Server: Based on the analysis results, the server generates corrected images that emphasize information that requires attention.

[0294] Server: If the emotion engine recognizes the user's surprise or fear, it prioritizes displaying information that is considered particularly dangerous.

[0295] Terminal: This information is projected onto the retina, allowing the user to respond appropriately.

[0296] In this way, the system effectively supports the vision of glaucoma patients and also supports a safe and comfortable life by providing appropriate information according to their emotions.

[0297] The processing flow will be explained below.

[0298] Step 1:

[0299] Device (camera): A camera mounted on glasses worn by the user captures images of the surroundings in real time, and these images are stored as digital data frame by frame.

[0300] Step 2:

[0301] Terminal: Captured video data is sent to the AI ​​analysis server via wireless communication, using a low-latency, highly reliable communication protocol.

[0302] Step 3:

[0303] Server: Stores the received video data in a buffer for analysis.

[0304] Step 4:

[0305] Server: The AI ​​model analyzes the video data stored in the buffer and recognizes objects in the image, such as cars, pedestrians, and traffic lights, and assigns labels and bounding boxes to each.

[0306] Step 5:

[0307] Server: Analyzes textual information contained in video using OCR technology and saves it as metadata. For example, extracts text from traffic lights and signs.

[0308] Step 6:

[0309] Server: Analyzes the user's emotional state in real time using an emotion engine. It uses facial expression recognition and voice analysis technology to detect emotional states (e.g., surprise, anxiety, joy, etc.).

[0310] Step 7:

[0311] Server: Based on the analysis results and the user's emotional state, an image is generated that corrects for the visual field defect. This involves enlarging parts of the image to make important information easier for the user to see.

[0312] Step 8:

[0313] Server: Further adjusts the correction content of the video data based on the user's emotional state recognized by the emotion engine. For example, if the user is feeling anxious, highlight warning information.

[0314] Step 9:

[0315] Server: Object information and extracted text information are overlaid on the original video. For example, a large "STOP" message is displayed over a "red light."

[0316] Step 10:

[0317] Server: Transmits video data optimized based on correction and emotion information to the terminal.

[0318] Step 11:

[0319] Terminal: The received corrected image data is projected onto the user's retina in real time, allowing the user to fill in information about the missing area of ​​their visual field.

[0320] Step 12:

[0321] User: Views the projected image and recognizes the surrounding situation and objects, enabling safe walking and driving.

[0322] Specific examples

[0323] When the user is waiting at a traffic light

[0324] 1. Device (camera): Captures the traffic light in front of the user.

[0325] 2. Terminal: Sends the captured video data to the server.

[0326] 3. Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, analyzes the text displayed on the traffic lights.

[0327] 4. Server: Generates corrected images based on the traffic light color and text information.

[0328] 5. Server: If the emotion engine recognizes the user's anxious state, it highlights the traffic light information.

[0329] 6. Server: Sends the corrected and highlighted traffic light information to the terminal.

[0330] 7. Terminal: The corrected image is projected onto the user's retina, allowing the user to see the signal with confidence.

[0331] If the user is walking down the street

[0332] 1. Terminal (camera): Captures images of the road and intersection ahead.

[0333] 2. Terminal: Sends video data to the server.

[0334] 3. Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[0335] 4. Server: Based on the analysis results, the server generates a corrected image by emphasizing information that requires attention.

[0336] 5. Server: If the emotion engine recognizes the user's surprise or fear, it prioritizes displaying information that is considered particularly dangerous.

[0337] 6. Server: Sends the corrected and enhanced information to the terminal.

[0338] 7. Terminal: This information is projected onto the retina, allowing the user to respond appropriately.

[0339] Through this series of steps, the visual assistance system effectively compensates for the user's visual field loss and also provides appropriate information according to their emotions, supporting a safe and comfortable life.

[0340] Example 2

[0341] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0342] Conventional visual assistance systems compensate for visual field defects in visually impaired people without taking the user's emotional state into consideration, which can result in information overload or overlooking important information. Furthermore, the provision of information adapted to the user's behavioral situation and surrounding environment is insufficient, making it difficult for the user to act safely and securely. This has led to a need for improved safety and comfort in everyday life.

[0343] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0344] In this invention, the server includes means for analyzing video data, means for correcting visual field defects, and means for analyzing the user's emotional state, thereby enabling the correction of video data and optimization of display content based on the user's emotional state.

[0345] A "camera" is a device for capturing images, and is built into a glasses-type device worn by a user.

[0346] An "AI-equipped server" is a computer system that receives and analyzes captured video data, and is a device equipped with an AI model that performs object and character recognition.

[0347] A "projection device" is a device for projecting corrected image data onto the user's retina in real time.

[0348] "Means for analyzing emotional state" refers to technology or devices for analyzing the user's facial expressions, voice, etc. to determine the user's emotional state in real time.

[0349] "Means for recognizing objects" refers to technology that identifies objects in video and assigns labels and bounding boxes to them.

[0350] "Means for extracting text information" refers to technology that analyzes text in video and obtains text data.

[0351] "Means to correct visual field defects" refers to technologies that enlarge or reduce images or overlay text information to compensate for the user's visual field defects.

[0352] "Means for generating video data in real time" refers to technology for instantly generating corrected video based on analyzed data.

[0353] The embodiment of this invention is a system that complements visual information for visually impaired users and provides optimal auxiliary information according to the user's emotional state. This system is composed of a glasses-type device at its core, a camera, an AI analysis server, an emotion engine, and a retinal projection device.

[0354] System Configuration

[0355] 1. Camera:

[0356] The camera is embedded in glasses worn by the user and captures images of the area in front of it in real time, capturing images at 30 frames per second and storing them as digital data.

[0357] 2. AI analysis server:

[0358] The server receives video data sent from the device and analyzes it using an AI model. Specifically, it uses deep learning models such as ResNet and YOLO to recognize objects in the video, assign labels and bounding boxes, and use OCR technology to extract text information from the video.

[0359] 3. Emotion Engine:

[0360] The server uses an emotion engine to analyze the user's emotional state in real time, which uses technology to analyze facial expressions and voice to determine the user's emotional state (surprise, anxiety, joy, etc.).

[0361] 4. Retinal projection device:

[0362] The device projects the corrected image data onto the user's retina in real time, allowing the user to compensate for the visual field defect and obtain optimal visual information according to their emotions.

[0363] Program processing details

[0364] Video capture and transmission:

[0365] When a user puts on the glasses, the camera activates and starts capturing images of the surrounding area, which are then sent to an AI analysis server via Wi-Fi, 4G, or 5G networks.

[0366] Video data analysis and correction:

[0367] The server analyzes the received video data, performs object and character recognition, and then makes corrections based on the analysis results, such as enlarging a portion of the video.

[0368] Emotional State Analysis:

[0369] The server uses an emotion engine to determine the user's emotional state and highlights information according to anxiety, surprise, etc.

[0370] Generate and project corrected image data:

[0371] The server generates the corrected image data and transmits it to the terminal, which then projects it into the user's field of vision via a retinal projection device.

[0372] Specific example of operation

[0373] 1. If the user is waiting at a traffic light:

[0374] The device (camera) captures images of traffic lights and sends them to the server.

[0375] The server analyzes the color and text displayed on the traffic light and generates a corrected image.

[0376] The server adjusts the display to highlight traffic light information because the emotion engine recognizes the user's anxiety.

[0377] The device projects highlighted traffic light information onto the user's retina, allowing them to see the traffic lights with confidence.

[0378] 2. If the user is walking down the street:

[0379] The device (camera) captures images of the road and intersection ahead and sends them to the server.

[0380] The server recognizes pedestrians, cars and traffic lights and determines their location and movement.

[0381] The server generates corrected images that highlight important information, and if the emotion engine recognizes surprise or fear, it prioritizes displaying dangerous information.

[0382] The device projects highlighted information onto the user's retina, helping them navigate safely.

[0383] Examples of prompts for generative AI models:

[0384] "If a user with impaired vision due to glaucoma or other conditions is waiting at a traffic light, the emotion engine should generate optimally corrected video data when it recognizes the user's state of anxiety."

[0385] In this way, a system is realized that provides optimal visual information to users and supports safe and comfortable daily life.

[0386] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0387] Step 1:

[0388] Device: The user wears the glasses-type device. The built-in camera automatically starts up and captures the image in front of the user in real time at a rate of 30 frames per second. The captured image is stored in the internal memory as digital data.

[0389] Input: Camera video (analog signal)

[0390] Output: Digital video data (real time)

[0391] Step 2:

[0392] Terminal: Captured video data is packaged frame by frame and sent to an AI analysis server via Wi-Fi, 4G, or 5G networks. Data is sent using a low-latency communication protocol and adaptive data transfer technology according to the transmission conditions.

[0393] Input: Digital video data (internal memory)

[0394] Output: Data packets sent over the network

[0395] Step 3:

[0396] Server: Receives the transmitted video data and loads it into a data queue. Then, it uses a deep learning model (e.g., ResNet, YOLO) to analyze objects in the video in real time. During the analysis process, objects are assigned labels and bounding boxes.

[0397] Input: Video data received via the network

[0398] Output: Object recognition data (labels, bounding boxes)

[0399] Step 4:

[0400] Server: Analyzes text information in video using OCR technology. Specifically, it extracts text data from traffic lights, road signs, etc.

[0401] Input: Analyzed video data

[0402] Output: Text data (e.g., "Watch out for pedestrians," "The traffic light is red")

[0403] Step 5:

[0404] Server: Based on the analysis results, the image of the visual field defect is corrected. Part of the image is enlarged to make important information easier to see, and necessary text information is overlaid. During this process, warnings and cautions are highlighted.

[0405] Input: object recognition data, text data

[0406] Output: Corrected video data

[0407] Step 6:

[0408] Server: Analyzes the user's emotional state using an emotion engine. It uses voice and facial expression data acquired from the device's built-in microphone and camera to determine the user's emotions (e.g., surprise, anxiety, joy).

[0409] Input: User's voice and facial expression data

[0410] Output: Emotional state data

[0411] Step 7:

[0412] Server: Integrates emotional state data and video analysis data to generate optimally corrected video data. If the user feels anxious, it generates video that highlights warning information.

[0413] Input: Corrected video data, emotional state data

[0414] Output: Optimized corrected video data

[0415] Step 8:

[0416] Server: Sends the generated optimized corrected image data to the terminal. The data is compressed and transmitted quickly and efficiently.

[0417] Input: Optimized corrected image data

[0418] Output: Data packets sent over the network

[0419] Step 9:

[0420] Terminal: The received corrected image data is projected onto the user's retina in real time. The retinal projection device visually presents information to compensate for the visual field defect, allowing the user to act safely and effectively.

[0421] Input: Received corrected video data

[0422] Output: Visual information projected onto the retina

[0423] (Application example 2)

[0424] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0425] Existing visual assistance systems are limited to supplementing visual information for users with visual impairments and do not take into account the user's emotional state. Furthermore, they lack the ability to link with autonomous vehicles and provide real-time safety information inside and outside the vehicle. This has resulted in insufficient safety and comfort for visually impaired people in certain environments.

[0426] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0427] In this invention, the server includes a means for capturing video using a camera, a means equipped with AI for analyzing the captured video, a means for correcting the analyzed video data and projecting it onto the user's retina, a means for linking with an emotion engine that recognizes the user's emotional state and adjusts the information display, and a means for linking with an autonomous vehicle and presenting important information inside and outside the vehicle. This makes it possible to provide optimal visual information according to the user's emotional state, and by linking with an autonomous vehicle, it is possible to improve the safety and comfort of travel for visually impaired people.

[0428] A "camera" is a device for capturing visual information, acquiring video and images as digital data.

[0429] An "AI-equipped server" is a server that uses artificial intelligence technology to analyze video data and has the ability to recognize objects and characters.

[0430] A "projection device" is a device for projecting corrected image data directly onto the user's visual organs (such as the retina).

[0431] The "emotion engine" is an engine that recognizes the user's emotional state and adjusts the operation and display content of other systems based on that information.

[0432] An "autonomous vehicle" is a vehicle that can navigate roads autonomously without the need for human operation.

[0433] "Object recognition" is a technology that detects objects in an image and identifies what they are.

[0434] "Means for extracting text information" refers to technology that analyzes and extracts text from video.

[0435] "Correction of visual field defects" is a technology that allows visually impaired users to complement the parts they cannot see and reconstruct images.

[0436] "Means for highlighting information" refers to a technique for visually highlighting and displaying important information.

[0437] "Means of generating data in real time" refers to technology that instantly analyzes and processes acquired data and immediately displays or outputs the results.

[0438] "Data integration" refers to technology that allows data to be shared between different systems and utilized mutually.

[0439] To realize this application example, a visual aid system is required. The core system configuration is as follows:

[0440] System configuration

[0441] 1. Camera

[0442] Overview: Visual information is captured in real time using a camera mounted on a device worn by the user (such as smart glasses).

[0443] Hardware used: Camera in smart glasses (e.g. Google Glass)

[0444] 2. AI-powered servers

[0445] Overview: Analyzes video data sent from a camera and performs object and character recognition. Artificial intelligence technology is used for the specific analysis.

[0446] Software and hardware used: AI analysis server (e.g., AWS EC2, Google Cloud AI)

[0447] 3. Projection device

[0448] Overview: A device that projects corrected image data onto the user's retina, compensating for the user's visual impairment.

[0449] Hardware used: Retinal projection device in smart glasses

[0450] 4. Emotion Engine

[0451] Overview: Analyzes the user's emotional state in real time and displays the most appropriate information based on that emotion. Uses facial expression recognition and voice analysis technology.

[0452] Software used: Sentiment analysis engine (e.g. Microsoft Azure Emotion API)

[0453] 5. Collaboration with autonomous vehicles

[0454] Overview: Connecting data to autonomous vehicles provides users with important information about the inside and outside of the vehicle.

[0455] Hardware and software used: Communication interfaces for autonomous vehicles

[0456] Specific use cases

[0457] Example 1: Traffic light recognition and highlighting

[0458] When the user approaches an intersection, the camera captures the traffic light ahead. This video data is sent to an AI analysis server, which recognizes the color and position of the traffic light. At the same time, the emotion engine analyzes the user's emotional state and highlights the traffic light information if the user is feeling anxious. A projection device projects this information onto the user's retina, allowing the user to see the highlighted traffic light information.

[0459] Example 2: Autonomous vehicle crossing a pedestrian crossing

[0460] When a user is in an autonomous vehicle, a camera captures video of the crosswalk ahead. This video data is sent to an AI analysis server, which recognizes the movements of pedestrians and vehicles on the crosswalk. At the same time, an emotion engine analyzes the user's emotional state and highlights information that is considered particularly dangerous if the user feels surprised. A projection device projects this information onto the user's retina, allowing the user to safely understand the situation at the crosswalk.

[0461] Example prompts to input to the generative AI model

[0462] "Please tell me a program that can analyze video data in real time on an AI analysis server and recognize specific objects."

[0463] "How can I use a sentiment analysis engine to detect a user's emotional state in real time?"

[0464] "Please tell me the specific steps for correcting image data with smart glasses and using a retinal projection device."

[0465] In this way, the invention can provide visually impaired users with optimal visual information according to their emotions, and can support safe and comfortable travel in cooperation with autonomous vehicles.

[0466] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0467] System processing steps

[0468] Step 1: Capture footage

[0469] The device uses a camera to capture visual information in real time. Specifically, a camera mounted on smart glasses worn by the user captures images of the scene ahead and stores them as digital data frame by frame.

[0470] Input: Front view of the user

[0471] Output: Captured video frame data

[0472] Step 2: Sending video data

[0473] The device transmits the captured video frame data to an AI analysis server via the Internet, enabling remote, high-performance analysis.

[0474] Input: Video frame data

[0475] Output: Send data to the AI ​​analysis server

[0476] Step 3: Analyzing the video data

[0477] The server processes the received video frame data in real time, using an AI model (e.g., the YOLO object recognition model, analysis using TensorFlow) to recognize objects and characters in the video, identify objects such as vehicles and traffic lights, and assign labels and bounding boxes to them.

[0478] Input: Transmitted video frame data

[0479] Output: Object and character recognition results (labels, bounding boxes)

[0480] Step 4: Recognizing your emotional state

[0481] The server uses an emotion engine to analyze the user's emotional state in real time, using voice analysis and facial expression recognition technology, e.g., Microsoft Azure Emotion API.

[0482] Input: User's voice and facial expression data

[0483] Output: Emotional state recognition result (surprise, anxiety, joy, etc.)

[0484] Step 5: Correcting the video data

[0485] The server generates an image that corrects for the visual field defect based on the analysis results, and processes the data by enlarging parts of the image to make important information easier for the user to see.

[0486] Input: Object recognition and character recognition results, emotional state recognition results

[0487] Output: Corrected video data

[0488] Step 6: Performing Highlighting

[0489] The server optimizes the video data correction and display content based on the user's emotional state detected by the emotion engine. For example, if the user is feeling anxious, important warning information will be highlighted.

[0490] Input: Corrected video data, emotional state recognition results

[0491] Output: Video data optimized for the user's emotional state

[0492] Step 7: Transmitting and projecting video data

[0493] The server then transmits the optimized image data based on the correction and emotion information to the terminal, which then projects the received corrected image data onto the user's retina in real time, providing the user with the necessary visual information.

[0494] Input: Optimized video data

[0495] Output: Visual information projected onto the retina

[0496] Example processing steps

[0497] Example 1: Traffic light recognition and highlighting

[0498] Step 1: The device captures the image of the traffic light ahead.

[0499] Input: Traffic light image

[0500] Output: Captured video frame data

[0501] Step 2: The device sends the captured video data to the AI ​​analysis server.

[0502] Input: Video frame data

[0503] Output: Send data to the AI ​​analysis server

[0504] Step 3: The server recognizes the traffic light as an object and identifies its color (red, yellow, green) and location.

[0505] Input: Transmitted video frame data

[0506] Output: Object recognition and traffic light color / position data

[0507] Step 4: The server recognizes the user's anxiety state using the emotion engine.

[0508] Input: User's voice and facial expression data

[0509] Output: Emotional state recognition result (anxiety)

[0510] Step 5: The server generates a corrected image by highlighting the traffic light information.

[0511] Input: Traffic light color / position data, emotional state recognition results

[0512] Output: Highlighted corrected video data

[0513] Step 6: The server sends the highlighted corrected video data to the terminal.

[0514] Input: Highlighted corrected video data

[0515] Output: Sending data to the terminal

[0516] Step 7: The device projects the highlighted corrected image data onto the user's retina.

[0517] Input: Highlighted corrected video data

[0518] Output: Visual information projected onto the retina

[0519] Example prompts to input to the generative AI model

[0520] "Please tell me a program that can analyze video data in real time on an AI analysis server and recognize specific objects."

[0521] "How can I use a sentiment analysis engine to detect a user's emotional state in real time?"

[0522] "Please tell me the specific steps for correcting image data with smart glasses and using a retinal projection device."

[0523] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0524] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0525] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0526] [Second embodiment]

[0527] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0528] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0529] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0530] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0531] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0532] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0533] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0534] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0535] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0536] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0537] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0538] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0539] An embodiment of this invention provides a system that complements visual information for users with impaired vision due to glaucoma, etc. This system is centered around a glasses-type device worn by the user, and is composed of a camera, an AI analysis server, and a retinal projection device.

[0540] System configuration

[0541] 1. Camera:

[0542] It is installed in glasses worn by the user and captures images in front of the vehicle in real time, which are then saved as image frame data.

[0543] 2. AI analysis server:

[0544] The system receives video data sent from the camera and analyzes it using an AI model. Specifically, it recognizes objects and characters in the video and identifies the information.

[0545] To compensate for the visual field defects, the analysis results are used to generate corrected video data, which includes scaling the video and adding text information.

[0546] 3. Retinal projection device:

[0547] This device projects corrected image data onto the user's retina in real time, allowing the user to compensate for visual field defects and easily obtain visual information in everyday life.

[0548] Program processing (natural language explanation)

[0549] Video capture and transmission

[0550] Device (camera): The camera in the glasses worn by the user captures the image in front of the user in real time. This image is saved as data frame by frame.

[0551] Terminal: Sends captured video data to an AI analysis server via the internet.

[0552] Video data analysis

[0553] Server: The AI ​​analytics server processes the received video data in real time, using AI models to recognize objects in the video, such as cars, park benches, and traffic lights, and assigns bounding boxes to each.

[0554] Server: Using OCR technology, character information in the video is extracted, obtaining text data such as "Watch out for pedestrians" or "The traffic light is red."

[0555] Video data correction

[0556] Server: Based on the analysis results, an image is generated to compensate for the visual field defect. Important information is reconstructed in a form that is easy for the user to recognize, using techniques such as image enlargement.

[0557] Server: Object information and extracted text information are overlaid on the video in a simple format.

[0558] Video data transmission and projection

[0559] Server: Sends the corrected video data to the user terminal.

[0560] Terminal: The transmitted corrected image data is projected onto the user's retina in real time. This process complements the missing information in the user's field of vision.

[0561] Specific examples

[0562] 1. If the user is waiting at a traffic light:

[0563] Terminal (camera): Captures the traffic light in front of the user.

[0564] Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, extracts the traffic light's text information (e.g., "pedestrian signal").

[0565] Server: Reconstructs the traffic light color and text information as a corrected image.

[0566] Device: Projects traffic light colors and text information onto the user's retina to complement visual information.

[0567] 2. If the user is walking down the street:

[0568] Terminal (camera): Captures images of the road and intersections ahead.

[0569] Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[0570] Server: Emphasizes information about pedestrians and vehicles that require attention and generates corrected images.

[0571] Terminal: This information is projected onto the user's retina to complement the visual information.

[0572] In this way, the system effectively supports the vision of glaucoma patients, helping them live safe and comfortable lives.

[0573] The processing flow will be explained below.

[0574] Step 1:

[0575] Device (camera): A camera mounted on glasses worn by the user captures images of the surroundings in real time, and these images are stored as digital data frame by frame.

[0576] Step 2:

[0577] Terminal: Captured video data is sent to the AI ​​analysis server via the internet using a low-latency, highly reliable communication protocol.

[0578] Step 3:

[0579] Server: Stores the received video data in a buffer for analysis.

[0580] Step 4:

[0581] Server: The AI ​​model analyzes the video data stored in the buffer and recognizes objects in the image, such as cars, pedestrians, and traffic lights, and assigns labels and bounding boxes to each.

[0582] Step 5:

[0583] Server: Analyzes text information contained in video using OCR technology and saves it as metadata. For example, extracts text from traffic lights and signs.

[0584] Step 6:

[0585] Server: Based on the analysis results, an image is generated that corrects for the visual field defect. Here, processing such as enlarging parts of the image is performed to make important information easier for the user to see.

[0586] Step 7:

[0587] Server: Object information and extracted text information are overlaid on the original video. For example, a large "STOP" message is displayed over a "red light."

[0588] Step 8:

[0589] Server: Converts the corrected video data into a format suitable for head-up displays and sends it to the user's device.

[0590] Step 9:

[0591] Terminal: The received corrected image data is projected onto the user's retina in real time, allowing the user to fill in information about the missing area of ​​their visual field.

[0592] Step 10:

[0593] User: Views the projected image and recognizes the surrounding situation and objects, enabling safe walking and driving.

[0594] Through this series of steps, the visual aid system effectively compensates for the user's visual field defects and provides the visual information necessary for daily life.

[0595] Example 1

[0596] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0597] Visually impaired users face challenges in moving safely and accurately grasping information about their surroundings in daily life. In particular, if they have lost vision due to diseases such as glaucoma, they are more likely to miss important visual information in the lost area, which could lead to serious accidents or danger.

[0598] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0599] In this invention, the server includes means for recognizing objects in the video, means for extracting text information, means for correcting the visual field defect and reconstructing the video, and means for enlarging the corrected video and adding text information, thereby enabling visually impaired users to supplement important information present in the visual field defect in real time and recognize their surroundings safely and reliably.

[0600] A "camera" is a photographic device for capturing images.

[0601] "Video data" refers to visual information captured by a camera and recorded in digital format.

[0602] "Artificial intelligence" is a technology in which computer programs mimic human intelligence to solve problems, and in this case it is used to analyze video data.

[0603] A "server" is a computer device that receives data over a network and analyzes and processes it.

[0604] A "projection device" is a device for projecting the analyzed and corrected image data onto the user's retina.

[0605] A "network" is an infrastructure for data communication, and is used here to transmit data between the camera and the server.

[0606] "Object recognition" is a technology that detects specific objects in video data and identifies their location and type.

[0607] "Text information" refers to characters and text data contained in video data.

[0608] "OCR" stands for Optical Character Recognition, a technology that extracts text information from video data.

[0609] The "visual field defect area" refers to an area that is missing from the user's field of vision due to a visual impairment.

[0610] "Image correction" is a correction process carried out to compensate for missing parts of the field of view based on analyzed image data.

[0611] "Real-time" refers to processing and display occurring immediately, without delay.

[0612] "Image enlargement" is a technique for displaying a specific image portion in a larger, more easily visible manner.

[0613] "Adding text information" is a technique for overlaying analyzed text information onto video.

[0614] This invention provides a system that complements visual information for users with visual impairments due to glaucoma, etc. The system consists of a camera, an AI analysis server, a network, and a projection device.

[0615] 1. Camera: The camera is mounted on a glasses-type device worn by the user and captures the image in front of the user in real time. The captured image is recorded as digital data frame by frame.

[0616] 2. AI analysis server: The video data captured by the camera is sent to the AI ​​analysis server via the network. The server performs the following processing steps:

[0617] Object Recognition: Using AI models to identify objects in a video (cars, traffic lights, pedestrians, etc.) and assign them bounding boxes, for example, identifying a traffic light as red.

[0618] Text information extraction: Extract text information (e.g., "pedestrian signal") from the video using OCR technology.

[0619] Image correction: To generate a corrected image, areas of vision defects are filled in and important information is emphasized, such as overlaying traffic light colors or important messages on the image.

[0620] 3. Network: Wi-Fi or 4G / 5G networks are used for data communication, sending and receiving data between the camera and the server.

[0621] 4. Projection device: Projects corrected image data onto the user's retina in real time, allowing the user to instantly see important information present in the area of ​​visual field loss.

[0622] Specific examples

[0623] 1. If the user is waiting at a traffic light:

[0624] Terminal (camera): Captures the traffic light in front of the user.

[0625] Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, extracts the traffic light's text information (e.g., "pedestrian signal").

[0626] Server: Reconstructs the traffic light color and text information as a corrected image.

[0627] Device: Projects traffic light colors and text information onto the user's retina to complement visual information.

[0628] 2. If the user is walking down the street:

[0629] Terminal (camera): Captures images of the road and intersections ahead.

[0630] Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[0631] Server: Emphasizes information about pedestrians and vehicles that require attention and generates corrected images.

[0632] Terminal: This information is projected onto the user's retina to complement the visual information.

[0633] Prompt Sentence Examples

[0634] Traffic light discrimination prompt: The prompt for extracting color and text information from traffic light images, generating corrected images for retinal projection is as follows: "Please distinguish between red, yellow, and green signals from traffic light images, extract the traffic light text information, and generate corrected images. Then, send the generated images to the retinal projection device."

[0635] Object Recognition Prompt: The prompt for recognizing cars, pedestrians, and traffic lights at intersections from the forward video while the user is walking and generating corrected video for safety information is as follows: "Please recognize cars, pedestrians, and traffic lights from the forward video while the user is walking, and reconstruct the corrected video for safety information. Then, send the generated video to the retinal projection device."

[0636] This system will enable glaucoma patients and other visually impaired users to effectively supplement their visual information in their daily lives, enabling them to live safe and comfortable lives.

[0637] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0638] Step 1:

[0639] The user turns on the eyeglasses-type device.

[0640] Input: Power-on operation

[0641] Specific operation: The user presses a button to turn on the device, which starts up the device's system and initializes the camera and retinal projection device.

[0642] Output: Device booted, camera and retinal projection device ready

[0643] Step 2:

[0644] The device (camera) captures the image in front of the user.

[0645] Input: Device boot complete, user visible

[0646] How it works: The camera is installed inside the device and captures real-time images of what is in front of the user's field of vision.

[0647] Output: Captured video data (frame by frame)

[0648] Step 3:

[0649] The device (camera) sends the captured video data to the server.

[0650] Input: Captured video data (frame by frame)

[0651] Specific operation: The captured video data is sent to an AI analysis server via Wi-Fi or 4G / 5G networks.

[0652] Output: Video data received by the server

[0653] Step 4:

[0654] The server inputs the received video data into the AI ​​model and begins analysis.

[0655] Input: Video data received by the server

[0656] How it works: The server inputs the video data into the AI ​​model and performs object recognition and OCR processing, such as identifying traffic lights, pedestrians, and text information and enclosing them in bounding boxes.

[0657] Output: Analyzed video data (object recognition information, text information)

[0658] Step 5:

[0659] The server generates an image that corrects the visual field defect based on the analysis results.

[0660] Input: Analyzed video data (object recognition information, text information)

[0661] How it works: The server uses the analysis results to generate an image that complements the visual field defect, enlarging important information and adding warning text such as "Watch your direction."

[0662] Output: Corrected video data

[0663] Step 6:

[0664] The server transmits the corrected video data to the user terminal.

[0665] Input: Corrected video data

[0666] Specific operation: The corrected video data is immediately sent to the user terminal and adjusted to minimize delay.

[0667] Output: Corrected video data received by the user device

[0668] Step 7:

[0669] The terminal (retinal projection device) projects the corrected image data onto the user's retina.

[0670] Input: Corrected video data received by the user terminal

[0671] How it works: The retinal projection device accurately projects the corrected image data into the user's field of vision, for example, displaying the red light of a traffic light and the text "Please stop" in the center of the field of vision.

[0672] Output: Corrected image projected onto the user's retina

[0673] In this way, each step of the system is performed sequentially, allowing visually impaired users to complete important visual information in real time.

[0674] (Application example 1)

[0675] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0676] This invention relates to a visual aid system that enables visually impaired users to appropriately acquire visual information and act safely in daily life and in specific situations. In particular, the objective is to provide a system that analyzes visual information in real time while driving a car and complements and emphasizes important information, thereby enabling visually impaired people to drive safely. Current visual aid systems lack analytical accuracy, real-time performance, and retinal projection technology, making improving driving safety a challenge.

[0677] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0678] In this invention, the server includes a means for capturing video using a camera, a computer equipped with artificial intelligence for analyzing the captured video, a projection device for correcting the analyzed video data and projecting it onto the user's retina, a means for transmitting the corrected video data from the computer to a terminal in real time, and a means for displaying the corrected video data on the terminal. This allows visually impaired users to accurately grasp visual information while driving, significantly improving safety. Furthermore, real-time analysis and visual enhancement allow users to immediately recognize important information such as obstacles and traffic signals, shortening the user's reaction time and making it easier to avoid danger.

[0679] A "camera" is a device for capturing images and is a piece of equipment used to record visual information.

[0680] "Artificial intelligence for video analysis" is a computer system that processes captured video data and performs object and character recognition.

[0681] A "computer" is a device that processes and analyzes data, and in particular, a computer equipped with artificial intelligence (AI)-based image analysis technology.

[0682] A "projection device" is a device that has the technology to project analyzed image data onto the user's vision, and in particular refers to equipment that projects images directly onto the retina.

[0683] "Real-time transmission means" refers to the technology and protocols for immediately transmitting processed data to the terminal.

[0684] A "terminal" is a device for displaying received video data, and includes smart glasses, head-mounted displays, and the like.

[0685] "Means for recognizing objects" refers to technology for detecting objects and people contained in video and identifying their location and type.

[0686] "Means for extracting text information" refers to OCR (optical character recognition) technology for analyzing and reading text data contained in video.

[0687] "Means for reconstructing images" refers to technology for correcting visual field defects and regenerating image data in a form that is easy for users to understand.

[0688] A "visual display device" is a device for visually presenting analyzed and corrected image data to a user, including smart glasses and retinal projection devices.

[0689] As an embodiment of the present invention, a system is provided for supplementing visual information for users with visual impairments such as glaucoma while driving a car. This system includes a camera, a computer (AI analysis server), a projection device, and a terminal. The role of each component and the overall processing flow are described in detail below.

[0690] System configuration

[0691] 1. Camera

[0692] The camera captures real-time images of the driving environment, which are then stored as data frame by frame and sent to a computer. The camera is used in smart glasses or head-mounted displays.

[0693] 2. AI analysis server

[0694] The computer (AI analysis server) receives the video data sent from the camera and analyzes it using an artificial intelligence model. Specifically, it recognizes objects in the video (cars, pedestrians, traffic lights, etc.) and extracts text information (intersection signs, road information, etc.). It then corrects for any visual defects and reconstructs the video in a way that emphasizes important information.

[0695] 3. Projection device

[0696] The corrected image data is projected directly onto the user's retina through a projection device, such as a device inside a pair of smart glasses, to complement the user's visual information in real time.

[0697] 4. Terminal

[0698] The corrected video data is sent in real time from the AI ​​analysis server to the device that displays the received video data, such as smart glasses or a head-mounted display.

[0699] Operational Overview

[0700] The specific steps involved in the operation of this system are described below.

[0701] 1. Video capture and transmission

[0702] The camera captures the image in front of the user and saves it as video frame data, which is then sent to an AI analysis server via the internet.

[0703] Example prompt sentence:

[0704] "Capture the video frame data to send to the AI ​​server and send it over the Internet. The video frame format is JPEG, and the server URL is 'http: / / ai-server / parse_frame'."

[0705] 2. Analysis and correction of video data

[0706] The server processes the received video data in real time. Using AI analysis, it recognizes objects in the video, identifying things like cars, park benches, and traffic lights. It also uses optical character recognition (OCR) technology to extract text information from the video. Based on the analysis results, it generates an image to complement any visual field defects, visually emphasizing important information.

[0707] Example prompt sentence:

[0708] "Please recognize objects and text contained in the received video data and return that information as a result. We use the YOLO model for object recognition and OCR technology for character recognition."

[0709] 3. Projecting the corrected image

[0710] The corrected image data sent from the server is received by the device and projected onto the user's retina in real time, allowing the user to supplement their visual information and accurately recognize important information while driving.

[0711] Example prompt sentence:

[0712] "Send the corrected image data to the retinal projection device so that the user can accurately perceive the visual information."

[0713] Specific examples

[0714] 1. Waiting at a traffic light

[0715] When the user approaches an intersection, the camera captures the traffic lights and intersection signs, and the AI ​​analysis server recognizes the color of the traffic lights (red, yellow, green) and provides that information to the user.

[0716] Example prompt sentence:

[0717] "Analyze the color and text information of traffic lights, and if the light is yellow, display 'Proceed with caution'."

[0718] 2. Walking on the street

[0719] As a user crosses the road, the camera captures the movements of pedestrians and vehicles. The AI ​​analytics server analyzes this information, highlights areas that require attention, and provides the user with enhanced video.

[0720] Example prompt sentence:

[0721] "Analyze the location of pedestrians and other vehicles as you approach an intersection and highlight any potentially dangerous situations."

[0722] This system will enable visually impaired people to obtain the visual information they need safely while driving a car, significantly improving driving safety.

[0723] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0724] Step 1:

[0725] A camera captures an image in front of the user.

[0726] Input: The user's current field of view.

[0727] Output: Video frame data.

[0728] How it works: The camera built into the smart glasses or head-mounted display captures video in real time and generates video frame data.

[0729] Step 2:

[0730] The device sends the captured video data to an AI analysis server.

[0731] Input: Video frame data.

[0732] Output: Video data sent to the server.

[0733] Specific operation: The device sends video data via the Internet, and the prompt to be entered is "Capture the video frame data to be sent to the AI ​​server and send it via the Internet. The video frame format is JPEG, and the server URL is 'http: / / ai-server / parse_frame'."

[0734] Step 3:

[0735] The server analyzes the received video data.

[0736] Input: Transmitted video data.

[0737] Output: Analysis results (object and character recognition data).

[0738] Specific operation: The server uses a generative AI model such as the YOLO model to recognize objects and text contained in the video data and identify information such as traffic lights, pedestrians, and signs. An example of a prompt is "Please recognize the objects and text contained in the received video data and return that information as a result. The YOLO model is used for object recognition, and OCR technology is used for character recognition."

[0739] Step 4:

[0740] The server corrects the video based on the analysis results.

[0741] Input: Analysis results.

[0742] Output: Corrected video data.

[0743] How it works: Based on the analysis results, the server compensates for the visual field defects and reconstructs the image to highlight important information. This correction makes the color of traffic lights and the content of signs appear larger, for example.

[0744] Step 5:

[0745] The server transmits the corrected video data to the terminal in real time.

[0746] Input: Corrected video data.

[0747] Output: Corrected video data sent to the device.

[0748] Specific operation: The server sends the corrected image data to the terminal in real time, and an example of a prompt is used: "Please send the corrected image data to the retinal projection device so that the user can accurately perceive the visual information."

[0749] Step 6:

[0750] The terminal displays the corrected image data on a display device (retinal projection device).

[0751] Input: Corrected video data.

[0752] Output: The corrected image displayed in the user's field of view.

[0753] How it works: The device sends the corrected image data to the retinal projection device, which projects it into the user's field of vision. For example, traffic light colors and sign information are projected directly onto the user's retina, allowing the user to perceive the necessary information, even if they have visual impairments.

[0754] By following the steps above, a visually impaired user can accurately obtain visual information while driving a car, enabling safe driving.

[0755] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0756] An embodiment of this invention provides a system that complements visual information for users with impaired vision due to glaucoma, etc., and also provides optimal auxiliary information according to the user's emotional state. This system is centered around a glasses-type device worn by the user, and is composed of a camera, an AI analysis server, an emotion engine, and a retinal projection device.

[0757] System configuration

[0758] 1. Camera:

[0759] The device is installed in glasses worn by the user and captures images of the surrounding area in real time, which are then stored as digital data frame by frame.

[0760] 2. AI analysis server:

[0761] The system receives video data sent from the camera and analyzes it using an AI model. Specifically, it recognizes objects and characters in the video and identifies the information.

[0762] To compensate for the visual field defects, the analysis results are used to generate corrected video data, which includes scaling the video and adding text information.

[0763] 3. Emotion Engine:

[0764] Recognize the user's emotional state in real time. The emotion engine can use technologies such as facial expression recognition and voice analysis.

[0765] Adjusting the display of visual information depending on emotional state (e.g., surprise, anxiety, joy, etc.).

[0766] 4. Retinal projection device:

[0767] This device projects corrected image data onto the user's retina in real time, allowing the user to compensate for visual field defects and easily obtain visual information in everyday life.

[0768] Program processing (natural language explanation)

[0769] Video capture and transmission

[0770] Device (camera): The camera in the glasses worn by the user captures the image in front of the user in real time, and this image is saved as digital data frame by frame.

[0771] Terminal: Sends captured video data to an AI analysis server via the internet.

[0772] Video data analysis

[0773] Server: The AI ​​analytics server processes the received video data in real time, using AI models to recognize objects in the video, such as cars, park benches, and traffic lights, and assigns labels and bounding boxes to each.

[0774] Server: Using OCR technology, the text information in the video is analyzed and text data such as "Watch out for pedestrians" or "The traffic light is red" is obtained.

[0775] Video data correction

[0776] Server: Based on the analysis results, an image is generated that corrects for the visual field defect. Here, processing such as enlarging parts of the image is performed to make important information easier for the user to see.

[0777] Server: Based on the object information and extracted text information, the server reconstructs the image with appropriate correction processing.

[0778] Emotion engine processing

[0779] Server: Analyzes the user's emotional state in real time using an emotion engine, for example, by using technology to detect emotions from facial expressions and voice.

[0780] Server: Based on the user's emotional state detected by the emotion engine, the server corrects the video data and optimizes the display content. For example, if the user is feeling anxious, it can highlight warning information.

[0781] Video data transmission and projection

[0782] Server: Transmits optimized video data based on correction and emotion information to the terminal.

[0783] Terminal: The received corrected image data is projected onto the user's retina in real time, allowing the user to compensate for the visual field defect and obtain appropriate visual information corresponding to their emotions.

[0784] Specific examples

[0785] 1. If the user is waiting at a traffic light:

[0786] Terminal (camera): Captures the traffic light in front of the user.

[0787] Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, analyzes the text displayed on the traffic lights.

[0788] Server: Generates corrected images based on the color and text information of traffic lights.

[0789] Server: If the emotion engine recognizes the user's state of anxiety, it highlights the traffic light information.

[0790] Device: Projects corrected and highlighted traffic light information onto the user's retina, allowing the user to see the traffic light with confidence.

[0791] 2. If the user is walking down the street:

[0792] Terminal (camera): Captures images of the road and intersections ahead.

[0793] Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[0794] Server: Based on the analysis results, the server generates corrected images that emphasize information that requires attention.

[0795] Server: If the emotion engine recognizes the user's surprise or fear, it prioritizes displaying information that is considered particularly dangerous.

[0796] Terminal: This information is projected onto the retina, allowing the user to respond appropriately.

[0797] In this way, the system effectively supports the vision of glaucoma patients and also supports a safe and comfortable life by providing appropriate information according to their emotions.

[0798] The processing flow will be explained below.

[0799] Step 1:

[0800] Device (camera): A camera mounted on glasses worn by the user captures images of the surroundings in real time, and these images are stored as digital data frame by frame.

[0801] Step 2:

[0802] Terminal: Captured video data is sent to the AI ​​analysis server via wireless communication, using a low-latency, highly reliable communication protocol.

[0803] Step 3:

[0804] Server: Stores the received video data in a buffer for analysis.

[0805] Step 4:

[0806] Server: The AI ​​model analyzes the video data stored in the buffer and recognizes objects in the image, such as cars, pedestrians, and traffic lights, and assigns labels and bounding boxes to each.

[0807] Step 5:

[0808] Server: Analyzes textual information contained in video using OCR technology and saves it as metadata. For example, extracts text from traffic lights and signs.

[0809] Step 6:

[0810] Server: Analyzes the user's emotional state in real time using an emotion engine. It uses facial expression recognition and voice analysis technology to detect emotional states (e.g., surprise, anxiety, joy, etc.).

[0811] Step 7:

[0812] Server: Based on the analysis results and the user's emotional state, an image is generated that corrects for the visual field defect. This involves enlarging parts of the image to make important information easier for the user to see.

[0813] Step 8:

[0814] Server: Further adjusts the correction content of the video data based on the user's emotional state recognized by the emotion engine. For example, if the user is feeling anxious, highlight warning information.

[0815] Step 9:

[0816] Server: Object information and extracted text information are overlaid on the original video. For example, a large "STOP" message is displayed over a "red light."

[0817] Step 10:

[0818] Server: Transmits video data optimized based on correction and emotion information to the terminal.

[0819] Step 11:

[0820] Terminal: The received corrected image data is projected onto the user's retina in real time, allowing the user to fill in information about the missing area of ​​their visual field.

[0821] Step 12:

[0822] User: Views the projected image and recognizes the surrounding situation and objects, enabling safe walking and driving.

[0823] Specific examples

[0824] When the user is waiting at a traffic light

[0825] 1. Device (camera): Captures the traffic light in front of the user.

[0826] 2. Terminal: Sends the captured video data to the server.

[0827] 3. Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, analyzes the text displayed on the traffic lights.

[0828] 4. Server: Generates corrected images based on the traffic light color and text information.

[0829] 5. Server: If the emotion engine recognizes the user's anxious state, it highlights the traffic light information.

[0830] 6. Server: Sends the corrected and highlighted traffic light information to the terminal.

[0831] 7. Terminal: The corrected image is projected onto the user's retina, allowing the user to see the signal with confidence.

[0832] If the user is walking down the street

[0833] 1. Terminal (camera): Captures images of the road and intersection ahead.

[0834] 2. Terminal: Sends video data to the server.

[0835] 3. Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[0836] 4. Server: Based on the analysis results, the server generates a corrected image by emphasizing information that requires attention.

[0837] 5. Server: If the emotion engine recognizes the user's surprise or fear, it prioritizes displaying information that is considered particularly dangerous.

[0838] 6. Server: Sends the corrected and enhanced information to the terminal.

[0839] 7. Terminal: This information is projected onto the retina, allowing the user to respond appropriately.

[0840] Through this series of steps, the visual assistance system effectively compensates for the user's visual field loss and also provides appropriate information according to their emotions, supporting a safe and comfortable life.

[0841] Example 2

[0842] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0843] Conventional visual assistance systems compensate for visual field defects in visually impaired people without taking the user's emotional state into consideration, which can result in information overload or overlooking important information. Furthermore, the provision of information adapted to the user's behavioral situation and surrounding environment is insufficient, making it difficult for the user to act safely and securely. This has led to a need for improved safety and comfort in everyday life.

[0844] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0845] In this invention, the server includes means for analyzing video data, means for correcting visual field defects, and means for analyzing the user's emotional state, thereby enabling the correction of video data and optimization of display content based on the user's emotional state.

[0846] A "camera" is a device for capturing images, and is built into a glasses-type device worn by a user.

[0847] An "AI-equipped server" is a computer system that receives and analyzes captured video data, and is a device equipped with an AI model that performs object and character recognition.

[0848] A "projection device" is a device for projecting corrected image data onto the user's retina in real time.

[0849] "Means for analyzing emotional state" refers to technology or devices for analyzing the user's facial expressions, voice, etc. to determine the user's emotional state in real time.

[0850] "Means for recognizing objects" refers to technology that identifies objects in video and assigns labels and bounding boxes to them.

[0851] "Means for extracting text information" refers to technology that analyzes text in video and obtains text data.

[0852] "Means to correct visual field defects" refers to technologies that enlarge or reduce images or overlay text information to compensate for the user's visual field defects.

[0853] "Means for generating video data in real time" refers to technology for instantly generating corrected video based on analyzed data.

[0854] The embodiment of this invention is a system that complements visual information for visually impaired users and provides optimal auxiliary information according to the user's emotional state. This system is composed of a glasses-type device at its core, a camera, an AI analysis server, an emotion engine, and a retinal projection device.

[0855] System Configuration

[0856] 1. Camera:

[0857] The camera is embedded in glasses worn by the user and captures images of the area in front of it in real time, capturing images at 30 frames per second and storing them as digital data.

[0858] 2. AI analysis server:

[0859] The server receives video data sent from the device and analyzes it using an AI model. Specifically, it uses deep learning models such as ResNet and YOLO to recognize objects in the video, assign labels and bounding boxes, and use OCR technology to extract text information from the video.

[0860] 3. Emotion Engine:

[0861] The server uses an emotion engine to analyze the user's emotional state in real time, which uses technology to analyze facial expressions and voice to determine the user's emotional state (surprise, anxiety, joy, etc.).

[0862] 4. Retinal projection device:

[0863] The device projects the corrected image data onto the user's retina in real time, allowing the user to compensate for the visual field defect and obtain optimal visual information according to their emotions.

[0864] Program processing details

[0865] Video capture and transmission:

[0866] When a user puts on the glasses, the camera activates and starts capturing images of the surrounding area, which are then sent to an AI analysis server via Wi-Fi, 4G, or 5G networks.

[0867] Video data analysis and correction:

[0868] The server analyzes the received video data, performs object and character recognition, and then makes corrections based on the analysis results, such as enlarging a portion of the video.

[0869] Emotional State Analysis:

[0870] The server uses an emotion engine to determine the user's emotional state and highlights information according to anxiety, surprise, etc.

[0871] Generate and project corrected image data:

[0872] The server generates the corrected image data and transmits it to the terminal, which then projects it into the user's field of vision via a retinal projection device.

[0873] Specific example of operation

[0874] 1. If the user is waiting at a traffic light:

[0875] The device (camera) captures images of traffic lights and sends them to the server.

[0876] The server analyzes the color and text displayed on the traffic light and generates a corrected image.

[0877] The server adjusts the display to highlight traffic light information because the emotion engine recognizes the user's anxiety.

[0878] The device projects highlighted traffic light information onto the user's retina, allowing them to see the traffic lights with confidence.

[0879] 2. If the user is walking down the street:

[0880] The device (camera) captures images of the road and intersection ahead and sends them to the server.

[0881] The server recognizes pedestrians, cars and traffic lights and determines their location and movement.

[0882] The server generates corrected images that highlight important information, and if the emotion engine recognizes surprise or fear, it prioritizes displaying dangerous information.

[0883] The device projects highlighted information onto the user's retina, helping them navigate safely.

[0884] Examples of prompts for generative AI models:

[0885] "If a user with impaired vision due to glaucoma or other conditions is waiting at a traffic light, the emotion engine should generate optimally corrected video data when it recognizes the user's state of anxiety."

[0886] In this way, a system is realized that provides optimal visual information to users and supports safe and comfortable daily life.

[0887] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0888] Step 1:

[0889] Device: The user wears the glasses-type device. The built-in camera automatically starts up and captures the image in front of the user in real time at a rate of 30 frames per second. The captured image is stored in the internal memory as digital data.

[0890] Input: Camera video (analog signal)

[0891] Output: Digital video data (real time)

[0892] Step 2:

[0893] Terminal: Captured video data is packaged frame by frame and sent to an AI analysis server via Wi-Fi, 4G, or 5G networks. Data is sent using a low-latency communication protocol and adaptive data transfer technology according to the transmission conditions.

[0894] Input: Digital video data (internal memory)

[0895] Output: Data packets sent over the network

[0896] Step 3:

[0897] Server: Receives the transmitted video data and loads it into a data queue. Then, it uses a deep learning model (e.g., ResNet, YOLO) to analyze objects in the video in real time. During the analysis process, objects are assigned labels and bounding boxes.

[0898] Input: Video data received via the network

[0899] Output: Object recognition data (labels, bounding boxes)

[0900] Step 4:

[0901] Server: Analyzes text information in video using OCR technology. Specifically, it extracts text data from traffic lights, road signs, etc.

[0902] Input: Analyzed video data

[0903] Output: Text data (e.g., "Watch out for pedestrians," "The traffic light is red")

[0904] Step 5:

[0905] Server: Based on the analysis results, the image of the visual field defect is corrected. Part of the image is enlarged to make important information easier to see, and necessary text information is overlaid. During this process, warnings and cautions are highlighted.

[0906] Input: object recognition data, text data

[0907] Output: Corrected video data

[0908] Step 6:

[0909] Server: Analyzes the user's emotional state using an emotion engine. It uses voice and facial expression data acquired from the device's built-in microphone and camera to determine the user's emotions (e.g., surprise, anxiety, joy).

[0910] Input: User's voice and facial expression data

[0911] Output: Emotional state data

[0912] Step 7:

[0913] Server: Integrates emotional state data and video analysis data to generate optimally corrected video data. If the user feels anxious, it generates video that highlights warning information.

[0914] Input: Corrected video data, emotional state data

[0915] Output: Optimized corrected video data

[0916] Step 8:

[0917] Server: Sends the generated optimized corrected image data to the terminal. The data is compressed and transmitted quickly and efficiently.

[0918] Input: Optimized corrected image data

[0919] Output: Data packets sent over the network

[0920] Step 9:

[0921] Terminal: The received corrected image data is projected onto the user's retina in real time. The retinal projection device visually presents information to compensate for the visual field defect, allowing the user to act safely and effectively.

[0922] Input: Received corrected video data

[0923] Output: Visual information projected onto the retina

[0924] (Application example 2)

[0925] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0926] Existing visual assistance systems are limited to supplementing visual information for users with visual impairments and do not take into account the user's emotional state. Furthermore, they lack the ability to link with autonomous vehicles and provide real-time safety information inside and outside the vehicle. This has resulted in insufficient safety and comfort for visually impaired people in certain environments.

[0927] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0928] In this invention, the server includes a means for capturing video using a camera, a means equipped with AI for analyzing the captured video, a means for correcting the analyzed video data and projecting it onto the user's retina, a means for linking with an emotion engine that recognizes the user's emotional state and adjusts the information display, and a means for linking with an autonomous vehicle and presenting important information inside and outside the vehicle. This makes it possible to provide optimal visual information according to the user's emotional state, and by linking with an autonomous vehicle, it is possible to improve the safety and comfort of travel for visually impaired people.

[0929] A "camera" is a device for capturing visual information, acquiring video and images as digital data.

[0930] An "AI-equipped server" is a server that uses artificial intelligence technology to analyze video data and has the ability to recognize objects and characters.

[0931] A "projection device" is a device for projecting corrected image data directly onto the user's visual organs (such as the retina).

[0932] The "emotion engine" is an engine that recognizes the user's emotional state and adjusts the operation and display content of other systems based on that information.

[0933] An "autonomous vehicle" is a vehicle that can navigate roads autonomously without the need for human operation.

[0934] "Object recognition" is a technology that detects objects in an image and identifies what they are.

[0935] "Means for extracting text information" refers to technology that analyzes and extracts text from video.

[0936] "Correction of visual field defects" is a technology that allows visually impaired users to complement the parts they cannot see and reconstruct images.

[0937] "Means for highlighting information" refers to a technique for visually highlighting and displaying important information.

[0938] "Means of generating data in real time" refers to technology that instantly analyzes and processes acquired data and immediately displays or outputs the results.

[0939] "Data integration" refers to technology that allows data to be shared between different systems and utilized mutually.

[0940] To realize this application example, a visual aid system is required. The core system configuration is as follows:

[0941] System configuration

[0942] 1. Camera

[0943] Overview: Visual information is captured in real time using a camera mounted on a device worn by the user (such as smart glasses).

[0944] Hardware used: Camera in smart glasses (e.g. Google Glass)

[0945] 2. AI-powered servers

[0946] Overview: Analyzes video data sent from a camera and performs object and character recognition. Artificial intelligence technology is used for the specific analysis.

[0947] Software and hardware used: AI analysis server (e.g., AWS EC2, Google Cloud AI)

[0948] 3. Projection device

[0949] Overview: A device that projects corrected image data onto the user's retina, compensating for the user's visual impairment.

[0950] Hardware used: Retinal projection device in smart glasses

[0951] 4. Emotion Engine

[0952] Overview: Analyzes the user's emotional state in real time and displays the most appropriate information based on that emotion. Uses facial expression recognition and voice analysis technology.

[0953] Software used: Sentiment analysis engine (e.g. Microsoft Azure Emotion API)

[0954] 5. Collaboration with autonomous vehicles

[0955] Overview: Connecting data to autonomous vehicles provides users with important information about the inside and outside of the vehicle.

[0956] Hardware and software used: Communication interfaces for autonomous vehicles

[0957] Specific use cases

[0958] Example 1: Traffic light recognition and highlighting

[0959] When the user approaches an intersection, the camera captures the traffic light ahead. This video data is sent to an AI analysis server, which recognizes the color and position of the traffic light. At the same time, the emotion engine analyzes the user's emotional state and highlights the traffic light information if the user is feeling anxious. A projection device projects this information onto the user's retina, allowing the user to see the highlighted traffic light information.

[0960] Example 2: Autonomous vehicle crossing a pedestrian crossing

[0961] When a user is in an autonomous vehicle, a camera captures video of the crosswalk ahead. This video data is sent to an AI analysis server, which recognizes the movements of pedestrians and vehicles on the crosswalk. At the same time, an emotion engine analyzes the user's emotional state and highlights information that is considered particularly dangerous if the user feels surprised. A projection device projects this information onto the user's retina, allowing the user to safely understand the situation at the crosswalk.

[0962] Example prompts to input to the generative AI model

[0963] "Please tell me a program that can analyze video data in real time on an AI analysis server and recognize specific objects."

[0964] "How can I use a sentiment analysis engine to detect a user's emotional state in real time?"

[0965] "Please tell me the specific steps for correcting image data with smart glasses and using a retinal projection device."

[0966] In this way, the invention can provide visually impaired users with optimal visual information according to their emotions, and can support safe and comfortable travel in cooperation with autonomous vehicles.

[0967] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0968] System processing steps

[0969] Step 1: Capture footage

[0970] The device uses a camera to capture visual information in real time. Specifically, a camera mounted on smart glasses worn by the user captures images of the scene ahead and stores them as digital data frame by frame.

[0971] Input: Front view of the user

[0972] Output: Captured video frame data

[0973] Step 2: Sending video data

[0974] The device transmits the captured video frame data to an AI analysis server via the Internet, enabling remote, high-performance analysis.

[0975] Input: Video frame data

[0976] Output: Send data to the AI ​​analysis server

[0977] Step 3: Analyzing the video data

[0978] The server processes the received video frame data in real time, using an AI model (e.g., the YOLO object recognition model, analysis using TensorFlow) to recognize objects and characters in the video, identify objects such as vehicles and traffic lights, and assign labels and bounding boxes to them.

[0979] Input: Transmitted video frame data

[0980] Output: Object and character recognition results (labels, bounding boxes)

[0981] Step 4: Recognizing your emotional state

[0982] The server uses an emotion engine to analyze the user's emotional state in real time, using voice analysis and facial expression recognition technology, e.g., Microsoft Azure Emotion API.

[0983] Input: User's voice and facial expression data

[0984] Output: Emotional state recognition result (surprise, anxiety, joy, etc.)

[0985] Step 5: Correcting the video data

[0986] The server generates an image that corrects for the visual field defect based on the analysis results, and processes the data by enlarging parts of the image to make important information easier for the user to see.

[0987] Input: Object recognition and character recognition results, emotional state recognition results

[0988] Output: Corrected video data

[0989] Step 6: Performing Highlighting

[0990] The server optimizes the video data correction and display content based on the user's emotional state detected by the emotion engine. For example, if the user is feeling anxious, important warning information will be highlighted.

[0991] Input: Corrected video data, emotional state recognition results

[0992] Output: Video data optimized for the user's emotional state

[0993] Step 7: Transmitting and projecting video data

[0994] The server then transmits the optimized image data based on the correction and emotion information to the terminal, which then projects the received corrected image data onto the user's retina in real time, providing the user with the necessary visual information.

[0995] Input: Optimized video data

[0996] Output: Visual information projected onto the retina

[0997] Example processing steps

[0998] Example 1: Traffic light recognition and highlighting

[0999] Step 1: The device captures the image of the traffic light ahead.

[1000] Input: Traffic light image

[1001] Output: Captured video frame data

[1002] Step 2: The device sends the captured video data to the AI ​​analysis server.

[1003] Input: Video frame data

[1004] Output: Send data to the AI ​​analysis server

[1005] Step 3: The server recognizes the traffic light as an object and identifies its color (red, yellow, green) and location.

[1006] Input: Transmitted video frame data

[1007] Output: Object recognition and traffic light color / position data

[1008] Step 4: The server recognizes the user's anxiety state using the emotion engine.

[1009] Input: User's voice and facial expression data

[1010] Output: Emotional state recognition result (anxiety)

[1011] Step 5: The server generates a corrected image by highlighting the traffic light information.

[1012] Input: Traffic light color / position data, emotional state recognition results

[1013] Output: Highlighted corrected video data

[1014] Step 6: The server sends the highlighted corrected video data to the terminal.

[1015] Input: Highlighted corrected video data

[1016] Output: Sending data to the terminal

[1017] Step 7: The device projects the highlighted corrected image data onto the user's retina.

[1018] Input: Highlighted corrected video data

[1019] Output: Visual information projected onto the retina

[1020] Example prompts to input to the generative AI model

[1021] "Please tell me a program that can analyze video data in real time on an AI analysis server and recognize specific objects."

[1022] "How can I use a sentiment analysis engine to detect a user's emotional state in real time?"

[1023] "Please tell me the specific steps for correcting image data with smart glasses and using a retinal projection device."

[1024] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1025] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1026] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1027] [Third embodiment]

[1028] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1029] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1031] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1032] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1033] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1035] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1036] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1038] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1039] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1040] An embodiment of this invention provides a system that complements visual information for users with impaired vision due to glaucoma, etc. This system is centered around a glasses-type device worn by the user, and is composed of a camera, an AI analysis server, and a retinal projection device.

[1041] System configuration

[1042] 1. Camera:

[1043] It is installed in glasses worn by the user and captures images in front of the vehicle in real time, which are then saved as image frame data.

[1044] 2. AI analysis server:

[1045] The system receives video data sent from the camera and analyzes it using an AI model. Specifically, it recognizes objects and characters in the video and identifies the information.

[1046] To compensate for the visual field defects, the analysis results are used to generate corrected video data, which includes scaling the video and adding text information.

[1047] 3. Retinal projection device:

[1048] This device projects corrected image data onto the user's retina in real time, allowing the user to compensate for visual field defects and easily obtain visual information in everyday life.

[1049] Program processing (natural language explanation)

[1050] Video capture and transmission

[1051] Device (camera): The camera in the glasses worn by the user captures the image in front of the user in real time. This image is saved as data frame by frame.

[1052] Terminal: Sends captured video data to an AI analysis server via the internet.

[1053] Video data analysis

[1054] Server: The AI ​​analytics server processes the received video data in real time, using AI models to recognize objects in the video, such as cars, park benches, and traffic lights, and assigns bounding boxes to each.

[1055] Server: Using OCR technology, character information in the video is extracted, obtaining text data such as "Watch out for pedestrians" or "The traffic light is red."

[1056] Video data correction

[1057] Server: Based on the analysis results, an image is generated to compensate for the visual field defect. Important information is reconstructed in a form that is easy for the user to recognize, using techniques such as image enlargement.

[1058] Server: Object information and extracted text information are overlaid on the video in a simple format.

[1059] Video data transmission and projection

[1060] Server: Sends the corrected video data to the user terminal.

[1061] Terminal: The transmitted corrected image data is projected onto the user's retina in real time. This process complements the missing information in the user's field of vision.

[1062] Specific examples

[1063] 1. If the user is waiting at a traffic light:

[1064] Terminal (camera): Captures the traffic light in front of the user.

[1065] Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, extracts the traffic light's text information (e.g., "pedestrian signal").

[1066] Server: Reconstructs the traffic light color and text information as a corrected image.

[1067] Device: Projects traffic light colors and text information onto the user's retina to complement visual information.

[1068] 2. If the user is walking down the street:

[1069] Terminal (camera): Captures images of the road and intersections ahead.

[1070] Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[1071] Server: Emphasizes information about pedestrians and vehicles that require attention and generates corrected images.

[1072] Terminal: This information is projected onto the user's retina to complement the visual information.

[1073] In this way, the system effectively supports the vision of glaucoma patients, helping them live safe and comfortable lives.

[1074] The processing flow will be explained below.

[1075] Step 1:

[1076] Device (camera): A camera mounted on glasses worn by the user captures images of the surroundings in real time, and these images are stored as digital data frame by frame.

[1077] Step 2:

[1078] Terminal: Captured video data is sent to the AI ​​analysis server via the internet using a low-latency, highly reliable communication protocol.

[1079] Step 3:

[1080] Server: Stores the received video data in a buffer for analysis.

[1081] Step 4:

[1082] Server: The AI ​​model analyzes the video data stored in the buffer and recognizes objects in the image, such as cars, pedestrians, and traffic lights, and assigns labels and bounding boxes to each.

[1083] Step 5:

[1084] Server: Analyzes text information contained in video using OCR technology and saves it as metadata. For example, extracts text from traffic lights and signs.

[1085] Step 6:

[1086] Server: Based on the analysis results, an image is generated that corrects for the visual field defect. Here, processing such as enlarging parts of the image is performed to make important information easier for the user to see.

[1087] Step 7:

[1088] Server: Object information and extracted text information are overlaid on the original video. For example, a large "STOP" message is displayed over a "red light."

[1089] Step 8:

[1090] Server: Converts the corrected video data into a format suitable for head-up displays and sends it to the user's device.

[1091] Step 9:

[1092] Terminal: The received corrected image data is projected onto the user's retina in real time, allowing the user to fill in information about the missing area of ​​their visual field.

[1093] Step 10:

[1094] User: Views the projected image and recognizes the surrounding situation and objects, enabling safe walking and driving.

[1095] Through this series of steps, the visual aid system effectively compensates for the user's visual field defects and provides the visual information necessary for daily life.

[1096] Example 1

[1097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1098] Visually impaired users face challenges in moving safely and accurately grasping information about their surroundings in daily life. In particular, if they have lost vision due to diseases such as glaucoma, they are more likely to miss important visual information in the lost area, which could lead to serious accidents or danger.

[1099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1100] In this invention, the server includes means for recognizing objects in the video, means for extracting text information, means for correcting the visual field defect and reconstructing the video, and means for enlarging the corrected video and adding text information, thereby enabling visually impaired users to supplement important information present in the visual field defect in real time and recognize their surroundings safely and reliably.

[1101] A "camera" is a photographic device for capturing images.

[1102] "Video data" refers to visual information captured by a camera and recorded in digital format.

[1103] "Artificial intelligence" is a technology in which computer programs mimic human intelligence to solve problems, and in this case it is used to analyze video data.

[1104] A "server" is a computer device that receives data over a network and analyzes and processes it.

[1105] A "projection device" is a device for projecting the analyzed and corrected image data onto the user's retina.

[1106] A "network" is an infrastructure for data communication, and is used here to transmit data between the camera and the server.

[1107] "Object recognition" is a technology that detects specific objects in video data and identifies their location and type.

[1108] "Text information" refers to characters and text data contained in video data.

[1109] "OCR" stands for Optical Character Recognition, a technology that extracts text information from video data.

[1110] The "visual field defect area" refers to an area that is missing from the user's field of vision due to a visual impairment.

[1111] "Image correction" is a correction process carried out to compensate for missing parts of the field of view based on analyzed image data.

[1112] "Real-time" refers to processing and display occurring immediately, without delay.

[1113] "Image enlargement" is a technique for displaying a specific image portion in a larger, more easily visible manner.

[1114] "Adding text information" is a technique for overlaying analyzed text information onto video.

[1115] This invention provides a system that complements visual information for users with visual impairments due to glaucoma, etc. The system consists of a camera, an AI analysis server, a network, and a projection device.

[1116] 1. Camera: The camera is mounted on a glasses-type device worn by the user and captures the image in front of the user in real time. The captured image is recorded as digital data frame by frame.

[1117] 2. AI analysis server: The video data captured by the camera is sent to the AI ​​analysis server via the network. The server performs the following processing steps:

[1118] Object Recognition: Using AI models to identify objects in a video (cars, traffic lights, pedestrians, etc.) and assign them bounding boxes, for example, identifying a traffic light as red.

[1119] Text information extraction: Extract text information (e.g., "pedestrian signal") from the video using OCR technology.

[1120] Image correction: To generate a corrected image, areas of vision defects are filled in and important information is emphasized, such as overlaying traffic light colors or important messages on the image.

[1121] 3. Network: Wi-Fi or 4G / 5G networks are used for data communication, sending and receiving data between the camera and the server.

[1122] 4. Projection device: Projects corrected image data onto the user's retina in real time, allowing the user to instantly see important information present in the area of ​​visual field loss.

[1123] Specific examples

[1124] 1. If the user is waiting at a traffic light:

[1125] Terminal (camera): Captures the traffic light in front of the user.

[1126] Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, extracts the traffic light's text information (e.g., "pedestrian signal").

[1127] Server: Reconstructs the traffic light color and text information as a corrected image.

[1128] Device: Projects traffic light colors and text information onto the user's retina to complement visual information.

[1129] 2. If the user is walking down the street:

[1130] Terminal (camera): Captures images of the road and intersections ahead.

[1131] Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[1132] Server: Emphasizes information about pedestrians and vehicles that require attention and generates corrected images.

[1133] Terminal: This information is projected onto the user's retina to complement the visual information.

[1134] Prompt Sentence Examples

[1135] Traffic light discrimination prompt: The prompt for extracting color and text information from traffic light images, generating corrected images for retinal projection is as follows: "Please distinguish between red, yellow, and green signals from traffic light images, extract the traffic light text information, and generate corrected images. Then, send the generated images to the retinal projection device."

[1136] Object Recognition Prompt: The prompt for recognizing cars, pedestrians, and traffic lights at intersections from the forward video while the user is walking and generating corrected video for safety information is as follows: "Please recognize cars, pedestrians, and traffic lights from the forward video while the user is walking, and reconstruct the corrected video for safety information. Then, send the generated video to the retinal projection device."

[1137] This system will enable glaucoma patients and other visually impaired users to effectively supplement their visual information in their daily lives, enabling them to live safe and comfortable lives.

[1138] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1139] Step 1:

[1140] The user turns on the eyeglasses-type device.

[1141] Input: Power-on operation

[1142] Specific operation: The user presses a button to turn on the device, which starts up the device's system and initializes the camera and retinal projection device.

[1143] Output: Device booted, camera and retinal projection device ready

[1144] Step 2:

[1145] The device (camera) captures the image in front of the user.

[1146] Input: Device boot complete, user visible

[1147] How it works: The camera is installed inside the device and captures real-time images of what is in front of the user's field of vision.

[1148] Output: Captured video data (frame by frame)

[1149] Step 3:

[1150] The device (camera) sends the captured video data to the server.

[1151] Input: Captured video data (frame by frame)

[1152] Specific operation: The captured video data is sent to an AI analysis server via Wi-Fi or 4G / 5G networks.

[1153] Output: Video data received by the server

[1154] Step 4:

[1155] The server inputs the received video data into the AI ​​model and begins analysis.

[1156] Input: Video data received by the server

[1157] How it works: The server inputs the video data into the AI ​​model and performs object recognition and OCR processing, such as identifying traffic lights, pedestrians, and text information and enclosing them in bounding boxes.

[1158] Output: Analyzed video data (object recognition information, text information)

[1159] Step 5:

[1160] The server generates an image that corrects the visual field defect based on the analysis results.

[1161] Input: Analyzed video data (object recognition information, text information)

[1162] How it works: The server uses the analysis results to generate an image that complements the visual field defect, enlarging important information and adding warning text such as "Watch your direction."

[1163] Output: Corrected video data

[1164] Step 6:

[1165] The server transmits the corrected video data to the user terminal.

[1166] Input: Corrected video data

[1167] Specific operation: The corrected video data is immediately sent to the user terminal and adjusted to minimize delay.

[1168] Output: Corrected video data received by the user device

[1169] Step 7:

[1170] The terminal (retinal projection device) projects the corrected image data onto the user's retina.

[1171] Input: Corrected video data received by the user terminal

[1172] How it works: The retinal projection device accurately projects the corrected image data into the user's field of vision, for example, displaying the red light of a traffic light and the text "Please stop" in the center of the field of vision.

[1173] Output: Corrected image projected onto the user's retina

[1174] In this way, each step of the system is performed sequentially, allowing visually impaired users to complete important visual information in real time.

[1175] (Application example 1)

[1176] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1177] This invention relates to a visual aid system that enables visually impaired users to appropriately acquire visual information and act safely in daily life and in specific situations. In particular, the objective is to provide a system that analyzes visual information in real time while driving a car and complements and emphasizes important information, thereby enabling visually impaired people to drive safely. Current visual aid systems lack analytical accuracy, real-time performance, and retinal projection technology, making improving driving safety a challenge.

[1178] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1179] In this invention, the server includes a means for capturing video using a camera, a computer equipped with artificial intelligence for analyzing the captured video, a projection device for correcting the analyzed video data and projecting it onto the user's retina, a means for transmitting the corrected video data from the computer to a terminal in real time, and a means for displaying the corrected video data on the terminal. This allows visually impaired users to accurately grasp visual information while driving, significantly improving safety. Furthermore, real-time analysis and visual enhancement allow users to immediately recognize important information such as obstacles and traffic signals, shortening the user's reaction time and making it easier to avoid danger.

[1180] A "camera" is a device for capturing images and is a piece of equipment used to record visual information.

[1181] "Artificial intelligence for video analysis" is a computer system that processes captured video data and performs object and character recognition.

[1182] A "computer" is a device that processes and analyzes data, and in particular, a computer equipped with artificial intelligence (AI)-based image analysis technology.

[1183] A "projection device" is a device that has the technology to project analyzed image data onto the user's vision, and in particular refers to equipment that projects images directly onto the retina.

[1184] "Real-time transmission means" refers to the technology and protocols for immediately transmitting processed data to the terminal.

[1185] A "terminal" is a device for displaying received video data, and includes smart glasses, head-mounted displays, and the like.

[1186] "Means for recognizing objects" refers to technology for detecting objects and people contained in video and identifying their location and type.

[1187] "Means for extracting text information" refers to OCR (optical character recognition) technology for analyzing and reading text data contained in video.

[1188] "Means for reconstructing images" refers to technology for correcting visual field defects and regenerating image data in a form that is easy for users to understand.

[1189] A "visual display device" is a device for visually presenting analyzed and corrected image data to a user, including smart glasses and retinal projection devices.

[1190] As an embodiment of the present invention, a system is provided for supplementing visual information for users with visual impairments such as glaucoma while driving a car. This system includes a camera, a computer (AI analysis server), a projection device, and a terminal. The role of each component and the overall processing flow are described in detail below.

[1191] System configuration

[1192] 1. Camera

[1193] The camera captures real-time images of the driving environment, which are then stored as data frame by frame and sent to a computer. The camera is used in smart glasses or head-mounted displays.

[1194] 2. AI analysis server

[1195] The computer (AI analysis server) receives the video data sent from the camera and analyzes it using an artificial intelligence model. Specifically, it recognizes objects in the video (cars, pedestrians, traffic lights, etc.) and extracts text information (intersection signs, road information, etc.). It then corrects for any visual defects and reconstructs the video in a way that emphasizes important information.

[1196] 3. Projection device

[1197] The corrected image data is projected directly onto the user's retina through a projection device, such as a device inside a pair of smart glasses, to complement the user's visual information in real time.

[1198] 4. Terminal

[1199] The corrected video data is sent in real time from the AI ​​analysis server to the device that displays the received video data, such as smart glasses or a head-mounted display.

[1200] Operational Overview

[1201] The specific steps involved in the operation of this system are described below.

[1202] 1. Video capture and transmission

[1203] The camera captures the image in front of the user and saves it as video frame data, which is then sent to an AI analysis server via the internet.

[1204] Example prompt sentence:

[1205] "Capture the video frame data to send to the AI ​​server and send it over the Internet. The video frame format is JPEG, and the server URL is 'http: / / ai-server / parse_frame'."

[1206] 2. Analysis and correction of video data

[1207] The server processes the received video data in real time. Using AI analysis, it recognizes objects in the video, identifying things like cars, park benches, and traffic lights. It also uses optical character recognition (OCR) technology to extract text information from the video. Based on the analysis results, it generates an image to complement any visual field defects, visually emphasizing important information.

[1208] Example prompt sentence:

[1209] "Please recognize objects and text contained in the received video data and return that information as a result. We use the YOLO model for object recognition and OCR technology for character recognition."

[1210] 3. Projecting the corrected image

[1211] The corrected image data sent from the server is received by the device and projected onto the user's retina in real time, allowing the user to supplement their visual information and accurately recognize important information while driving.

[1212] Example prompt sentence:

[1213] "Send the corrected image data to the retinal projection device so that the user can accurately perceive the visual information."

[1214] Specific examples

[1215] 1. Waiting at a traffic light

[1216] When the user approaches an intersection, the camera captures the traffic lights and intersection signs, and the AI ​​analysis server recognizes the color of the traffic lights (red, yellow, green) and provides that information to the user.

[1217] Example prompt sentence:

[1218] "Analyze the color and text information of traffic lights, and if the light is yellow, display 'Proceed with caution'."

[1219] 2. Walking on the street

[1220] As a user crosses the road, the camera captures the movements of pedestrians and vehicles. The AI ​​analytics server analyzes this information, highlights areas that require attention, and provides the user with enhanced video.

[1221] Example prompt sentence:

[1222] "Analyze the location of pedestrians and other vehicles as you approach an intersection and highlight any potentially dangerous situations."

[1223] This system will enable visually impaired people to obtain the visual information they need safely while driving a car, significantly improving driving safety.

[1224] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1225] Step 1:

[1226] A camera captures an image in front of the user.

[1227] Input: The user's current field of view.

[1228] Output: Video frame data.

[1229] How it works: The camera built into the smart glasses or head-mounted display captures video in real time and generates video frame data.

[1230] Step 2:

[1231] The device sends the captured video data to an AI analysis server.

[1232] Input: Video frame data.

[1233] Output: Video data sent to the server.

[1234] Specific operation: The device sends video data via the Internet, and the prompt to be entered is "Capture the video frame data to be sent to the AI ​​server and send it via the Internet. The video frame format is JPEG, and the server URL is 'http: / / ai-server / parse_frame'."

[1235] Step 3:

[1236] The server analyzes the received video data.

[1237] Input: Transmitted video data.

[1238] Output: Analysis results (object and character recognition data).

[1239] Specific operation: The server uses a generative AI model such as the YOLO model to recognize objects and text contained in the video data and identify information such as traffic lights, pedestrians, and signs. An example of a prompt is "Please recognize the objects and text contained in the received video data and return that information as a result. The YOLO model is used for object recognition, and OCR technology is used for character recognition."

[1240] Step 4:

[1241] The server corrects the video based on the analysis results.

[1242] Input: Analysis results.

[1243] Output: Corrected video data.

[1244] How it works: Based on the analysis results, the server compensates for the visual field defects and reconstructs the image to highlight important information. This correction makes the color of traffic lights and the content of signs appear larger, for example.

[1245] Step 5:

[1246] The server transmits the corrected video data to the terminal in real time.

[1247] Input: Corrected video data.

[1248] Output: Corrected video data sent to the device.

[1249] Specific operation: The server sends the corrected image data to the terminal in real time, and an example of a prompt is used: "Please send the corrected image data to the retinal projection device so that the user can accurately perceive the visual information."

[1250] Step 6:

[1251] The terminal displays the corrected image data on a display device (retinal projection device).

[1252] Input: Corrected video data.

[1253] Output: The corrected image displayed in the user's field of view.

[1254] How it works: The device sends the corrected image data to the retinal projection device, which projects it into the user's field of vision. For example, traffic light colors and sign information are projected directly onto the user's retina, allowing the user to perceive the necessary information, even if they have visual impairments.

[1255] By following the steps above, a visually impaired user can accurately obtain visual information while driving a car, enabling safe driving.

[1256] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1257] An embodiment of this invention provides a system that complements visual information for users with impaired vision due to glaucoma, etc., and also provides optimal auxiliary information according to the user's emotional state. This system is centered around a glasses-type device worn by the user, and is composed of a camera, an AI analysis server, an emotion engine, and a retinal projection device.

[1258] System configuration

[1259] 1. Camera:

[1260] The device is installed in glasses worn by the user and captures images of the surrounding area in real time, which are then stored as digital data frame by frame.

[1261] 2. AI analysis server:

[1262] The system receives video data sent from the camera and analyzes it using an AI model. Specifically, it recognizes objects and characters in the video and identifies the information.

[1263] To compensate for the visual field defects, the analysis results are used to generate corrected video data, which includes scaling the video and adding text information.

[1264] 3. Emotion Engine:

[1265] Recognize the user's emotional state in real time. The emotion engine can use technologies such as facial expression recognition and voice analysis.

[1266] Adjusting the display of visual information depending on emotional state (e.g., surprise, anxiety, joy, etc.).

[1267] 4. Retinal projection device:

[1268] This device projects corrected image data onto the user's retina in real time, allowing the user to compensate for visual field defects and easily obtain visual information in everyday life.

[1269] Program processing (natural language explanation)

[1270] Video capture and transmission

[1271] Device (camera): The camera in the glasses worn by the user captures the image in front of the user in real time, and this image is saved as digital data frame by frame.

[1272] Terminal: Sends captured video data to an AI analysis server via the internet.

[1273] Video data analysis

[1274] Server: The AI ​​analytics server processes the received video data in real time, using AI models to recognize objects in the video, such as cars, park benches, and traffic lights, and assigns labels and bounding boxes to each.

[1275] Server: Using OCR technology, the text information in the video is analyzed and text data such as "Watch out for pedestrians" or "The traffic light is red" is obtained.

[1276] Video data correction

[1277] Server: Based on the analysis results, an image is generated that corrects for the visual field defect. Here, processing such as enlarging parts of the image is performed to make important information easier for the user to see.

[1278] Server: Based on the object information and extracted text information, the server reconstructs the image with appropriate correction processing.

[1279] Emotion engine processing

[1280] Server: Analyzes the user's emotional state in real time using an emotion engine, for example, by using technology to detect emotions from facial expressions and voice.

[1281] Server: Based on the user's emotional state detected by the emotion engine, the server corrects the video data and optimizes the display content. For example, if the user is feeling anxious, it can highlight warning information.

[1282] Video data transmission and projection

[1283] Server: Transmits optimized video data based on correction and emotion information to the terminal.

[1284] Terminal: The received corrected image data is projected onto the user's retina in real time, allowing the user to compensate for the visual field defect and obtain appropriate visual information corresponding to their emotions.

[1285] Specific examples

[1286] 1. If the user is waiting at a traffic light:

[1287] Terminal (camera): Captures the traffic light in front of the user.

[1288] Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, analyzes the text displayed on the traffic lights.

[1289] Server: Generates corrected images based on the color and text information of traffic lights.

[1290] Server: If the emotion engine recognizes the user's state of anxiety, it highlights the traffic light information.

[1291] Device: Projects corrected and highlighted traffic light information onto the user's retina, allowing the user to see the traffic light with confidence.

[1292] 2. If the user is walking down the street:

[1293] Terminal (camera): Captures images of the road and intersections ahead.

[1294] Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[1295] Server: Based on the analysis results, the server generates corrected images that emphasize information that requires attention.

[1296] Server: If the emotion engine recognizes the user's surprise or fear, it prioritizes displaying information that is considered particularly dangerous.

[1297] Terminal: This information is projected onto the retina, allowing the user to respond appropriately.

[1298] In this way, the system effectively supports the vision of glaucoma patients and also supports a safe and comfortable life by providing appropriate information according to their emotions.

[1299] The processing flow will be explained below.

[1300] Step 1:

[1301] Device (camera): A camera mounted on glasses worn by the user captures images of the surroundings in real time, and these images are stored as digital data frame by frame.

[1302] Step 2:

[1303] Terminal: Captured video data is sent to the AI ​​analysis server via wireless communication, using a low-latency, highly reliable communication protocol.

[1304] Step 3:

[1305] Server: Stores the received video data in a buffer for analysis.

[1306] Step 4:

[1307] Server: The AI ​​model analyzes the video data stored in the buffer and recognizes objects in the image, such as cars, pedestrians, and traffic lights, and assigns labels and bounding boxes to each.

[1308] Step 5:

[1309] Server: Analyzes textual information contained in video using OCR technology and saves it as metadata. For example, extracts text from traffic lights and signs.

[1310] Step 6:

[1311] Server: Analyzes the user's emotional state in real time using an emotion engine. It uses facial expression recognition and voice analysis technology to detect emotional states (e.g., surprise, anxiety, joy, etc.).

[1312] Step 7:

[1313] Server: Based on the analysis results and the user's emotional state, an image is generated that corrects for the visual field defect. This involves enlarging parts of the image to make important information easier for the user to see.

[1314] Step 8:

[1315] Server: Further adjusts the correction content of the video data based on the user's emotional state recognized by the emotion engine. For example, if the user is feeling anxious, highlight warning information.

[1316] Step 9:

[1317] Server: Object information and extracted text information are overlaid on the original video. For example, a large "STOP" message is displayed over a "red light."

[1318] Step 10:

[1319] Server: Transmits video data optimized based on correction and emotion information to the terminal.

[1320] Step 11:

[1321] Terminal: The received corrected image data is projected onto the user's retina in real time, allowing the user to fill in information about the missing area of ​​their visual field.

[1322] Step 12:

[1323] User: Views the projected image and recognizes the surrounding situation and objects, enabling safe walking and driving.

[1324] Specific examples

[1325] When the user is waiting at a traffic light

[1326] 1. Device (camera): Captures the traffic light in front of the user.

[1327] 2. Terminal: Sends the captured video data to the server.

[1328] 3. Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, analyzes the text displayed on the traffic lights.

[1329] 4. Server: Generates corrected images based on the traffic light color and text information.

[1330] 5. Server: If the emotion engine recognizes the user's anxious state, it highlights the traffic light information.

[1331] 6. Server: Sends the corrected and highlighted traffic light information to the terminal.

[1332] 7. Terminal: The corrected image is projected onto the user's retina, allowing the user to see the signal with confidence.

[1333] If the user is walking down the street

[1334] 1. Terminal (camera): Captures images of the road and intersection ahead.

[1335] 2. Terminal: Sends video data to the server.

[1336] 3. Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[1337] 4. Server: Based on the analysis results, the server generates a corrected image by emphasizing information that requires attention.

[1338] 5. Server: If the emotion engine recognizes the user's surprise or fear, it prioritizes displaying information that is considered particularly dangerous.

[1339] 6. Server: Sends the corrected and enhanced information to the terminal.

[1340] 7. Terminal: This information is projected onto the retina, allowing the user to respond appropriately.

[1341] Through this series of steps, the visual assistance system effectively compensates for the user's visual field loss and also provides appropriate information according to their emotions, supporting a safe and comfortable life.

[1342] Example 2

[1343] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1344] Conventional visual assistance systems compensate for visual field defects in visually impaired people without taking the user's emotional state into consideration, which can result in information overload or overlooking important information. Furthermore, the provision of information adapted to the user's behavioral situation and surrounding environment is insufficient, making it difficult for the user to act safely and securely. This has led to a need for improved safety and comfort in everyday life.

[1345] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1346] In this invention, the server includes means for analyzing video data, means for correcting visual field defects, and means for analyzing the user's emotional state, thereby enabling the correction of video data and optimization of display content based on the user's emotional state.

[1347] A "camera" is a device for capturing images, and is built into a glasses-type device worn by a user.

[1348] An "AI-equipped server" is a computer system that receives and analyzes captured video data, and is a device equipped with an AI model that performs object and character recognition.

[1349] A "projection device" is a device for projecting corrected image data onto the user's retina in real time.

[1350] "Means for analyzing emotional state" refers to technology or devices for analyzing the user's facial expressions, voice, etc. to determine the user's emotional state in real time.

[1351] "Means for recognizing objects" refers to technology that identifies objects in video and assigns labels and bounding boxes to them.

[1352] "Means for extracting text information" refers to technology that analyzes text in video and obtains text data.

[1353] "Means to correct visual field defects" refers to technologies that enlarge or reduce images or overlay text information to compensate for the user's visual field defects.

[1354] "Means for generating video data in real time" refers to technology for instantly generating corrected video based on analyzed data.

[1355] The embodiment of this invention is a system that complements visual information for visually impaired users and provides optimal auxiliary information according to the user's emotional state. This system is composed of a glasses-type device at its core, a camera, an AI analysis server, an emotion engine, and a retinal projection device.

[1356] System Configuration

[1357] 1. Camera:

[1358] The camera is embedded in glasses worn by the user and captures images of the area in front of it in real time, capturing images at 30 frames per second and storing them as digital data.

[1359] 2. AI analysis server:

[1360] The server receives video data sent from the device and analyzes it using an AI model. Specifically, it uses deep learning models such as ResNet and YOLO to recognize objects in the video, assign labels and bounding boxes, and use OCR technology to extract text information from the video.

[1361] 3. Emotion Engine:

[1362] The server uses an emotion engine to analyze the user's emotional state in real time, which uses technology to analyze facial expressions and voice to determine the user's emotional state (surprise, anxiety, joy, etc.).

[1363] 4. Retinal projection device:

[1364] The device projects the corrected image data onto the user's retina in real time, allowing the user to compensate for the visual field defect and obtain optimal visual information according to their emotions.

[1365] Program processing details

[1366] Video capture and transmission:

[1367] When a user puts on the glasses, the camera activates and starts capturing images of the surrounding area, which are then sent to an AI analysis server via Wi-Fi, 4G, or 5G networks.

[1368] Video data analysis and correction:

[1369] The server analyzes the received video data, performs object and character recognition, and then makes corrections based on the analysis results, such as enlarging a portion of the video.

[1370] Emotional State Analysis:

[1371] The server uses an emotion engine to determine the user's emotional state and highlights information according to anxiety, surprise, etc.

[1372] Generate and project corrected image data:

[1373] The server generates the corrected image data and transmits it to the terminal, which then projects it into the user's field of vision via a retinal projection device.

[1374] Specific example of operation

[1375] 1. If the user is waiting at a traffic light:

[1376] The device (camera) captures images of traffic lights and sends them to the server.

[1377] The server analyzes the color and text displayed on the traffic light and generates a corrected image.

[1378] The server adjusts the display to highlight traffic light information because the emotion engine recognizes the user's anxiety.

[1379] The device projects highlighted traffic light information onto the user's retina, allowing them to see the traffic lights with confidence.

[1380] 2. If the user is walking down the street:

[1381] The device (camera) captures images of the road and intersection ahead and sends them to the server.

[1382] The server recognizes pedestrians, cars and traffic lights and determines their location and movement.

[1383] The server generates corrected images that highlight important information, and if the emotion engine recognizes surprise or fear, it prioritizes displaying dangerous information.

[1384] The device projects highlighted information onto the user's retina, helping them navigate safely.

[1385] Examples of prompts for generative AI models:

[1386] "If a user with impaired vision due to glaucoma or other conditions is waiting at a traffic light, the emotion engine should generate optimally corrected video data when it recognizes the user's state of anxiety."

[1387] In this way, a system is realized that provides optimal visual information to users and supports safe and comfortable daily life.

[1388] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1389] Step 1:

[1390] Device: The user wears the glasses-type device. The built-in camera automatically starts up and captures the image in front of the user in real time at a rate of 30 frames per second. The captured image is stored in the internal memory as digital data.

[1391] Input: Camera video (analog signal)

[1392] Output: Digital video data (real time)

[1393] Step 2:

[1394] Terminal: Captured video data is packaged frame by frame and sent to an AI analysis server via Wi-Fi, 4G, or 5G networks. Data is sent using a low-latency communication protocol and adaptive data transfer technology according to the transmission conditions.

[1395] Input: Digital video data (internal memory)

[1396] Output: Data packets sent over the network

[1397] Step 3:

[1398] Server: Receives the transmitted video data and loads it into a data queue. Then, it uses a deep learning model (e.g., ResNet, YOLO) to analyze objects in the video in real time. During the analysis process, objects are assigned labels and bounding boxes.

[1399] Input: Video data received via the network

[1400] Output: Object recognition data (labels, bounding boxes)

[1401] Step 4:

[1402] Server: Analyzes text information in video using OCR technology. Specifically, it extracts text data from traffic lights, road signs, etc.

[1403] Input: Analyzed video data

[1404] Output: Text data (e.g., "Watch out for pedestrians," "The traffic light is red")

[1405] Step 5:

[1406] Server: Based on the analysis results, the image of the visual field defect is corrected. Part of the image is enlarged to make important information easier to see, and necessary text information is overlaid. During this process, warnings and cautions are highlighted.

[1407] Input: object recognition data, text data

[1408] Output: Corrected video data

[1409] Step 6:

[1410] Server: Analyzes the user's emotional state using an emotion engine. It uses voice and facial expression data acquired from the device's built-in microphone and camera to determine the user's emotions (e.g., surprise, anxiety, joy).

[1411] Input: User's voice and facial expression data

[1412] Output: Emotional state data

[1413] Step 7:

[1414] Server: Integrates emotional state data and video analysis data to generate optimally corrected video data. If the user feels anxious, it generates video that highlights warning information.

[1415] Input: Corrected video data, emotional state data

[1416] Output: Optimized corrected video data

[1417] Step 8:

[1418] Server: Sends the generated optimized corrected image data to the terminal. The data is compressed and transmitted quickly and efficiently.

[1419] Input: Optimized corrected image data

[1420] Output: Data packets sent over the network

[1421] Step 9:

[1422] Terminal: The received corrected image data is projected onto the user's retina in real time. The retinal projection device visually presents information to compensate for the visual field defect, allowing the user to act safely and effectively.

[1423] Input: Received corrected video data

[1424] Output: Visual information projected onto the retina

[1425] (Application example 2)

[1426] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1427] Existing visual assistance systems are limited to supplementing visual information for users with visual impairments and do not take into account the user's emotional state. Furthermore, they lack the ability to link with autonomous vehicles and provide real-time safety information inside and outside the vehicle. This has resulted in insufficient safety and comfort for visually impaired people in certain environments.

[1428] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1429] In this invention, the server includes a means for capturing video using a camera, a means equipped with AI for analyzing the captured video, a means for correcting the analyzed video data and projecting it onto the user's retina, a means for linking with an emotion engine that recognizes the user's emotional state and adjusts the information display, and a means for linking with an autonomous vehicle and presenting important information inside and outside the vehicle. This makes it possible to provide optimal visual information according to the user's emotional state, and by linking with an autonomous vehicle, it is possible to improve the safety and comfort of travel for visually impaired people.

[1430] A "camera" is a device for capturing visual information, acquiring video and images as digital data.

[1431] An "AI-equipped server" is a server that uses artificial intelligence technology to analyze video data and has the ability to recognize objects and characters.

[1432] A "projection device" is a device for projecting corrected image data directly onto the user's visual organs (such as the retina).

[1433] The "emotion engine" is an engine that recognizes the user's emotional state and adjusts the operation and display content of other systems based on that information.

[1434] An "autonomous vehicle" is a vehicle that can navigate roads autonomously without the need for human operation.

[1435] "Object recognition" is a technology that detects objects in an image and identifies what they are.

[1436] "Means for extracting text information" refers to technology that analyzes and extracts text from video.

[1437] "Correction of visual field defects" is a technology that allows visually impaired users to complement the parts they cannot see and reconstruct images.

[1438] "Means for highlighting information" refers to a technique for visually highlighting and displaying important information.

[1439] "Means of generating data in real time" refers to technology that instantly analyzes and processes acquired data and immediately displays or outputs the results.

[1440] "Data integration" refers to technology that allows data to be shared between different systems and utilized mutually.

[1441] To realize this application example, a visual aid system is required. The core system configuration is as follows:

[1442] System configuration

[1443] 1. Camera

[1444] Overview: Visual information is captured in real time using a camera mounted on a device worn by the user (such as smart glasses).

[1445] Hardware used: Camera in smart glasses (e.g. Google Glass)

[1446] 2. AI-powered servers

[1447] Overview: Analyzes video data sent from a camera and performs object and character recognition. Artificial intelligence technology is used for the specific analysis.

[1448] Software and hardware used: AI analysis server (e.g., AWS EC2, Google Cloud AI)

[1449] 3. Projection device

[1450] Overview: A device that projects corrected image data onto the user's retina, compensating for the user's visual impairment.

[1451] Hardware used: Retinal projection device in smart glasses

[1452] 4. Emotion Engine

[1453] Overview: Analyzes the user's emotional state in real time and displays the most appropriate information based on that emotion. Uses facial expression recognition and voice analysis technology.

[1454] Software used: Sentiment analysis engine (e.g. Microsoft Azure Emotion API)

[1455] 5. Collaboration with autonomous vehicles

[1456] Overview: Connecting data to autonomous vehicles provides users with important information about the inside and outside of the vehicle.

[1457] Hardware and software used: Communication interfaces for autonomous vehicles

[1458] Specific use cases

[1459] Example 1: Traffic light recognition and highlighting

[1460] When the user approaches an intersection, the camera captures the traffic light ahead. This video data is sent to an AI analysis server, which recognizes the color and position of the traffic light. At the same time, the emotion engine analyzes the user's emotional state and highlights the traffic light information if the user is feeling anxious. A projection device projects this information onto the user's retina, allowing the user to see the highlighted traffic light information.

[1461] Example 2: Autonomous vehicle crossing a pedestrian crossing

[1462] When a user is in an autonomous vehicle, a camera captures video of the crosswalk ahead. This video data is sent to an AI analysis server, which recognizes the movements of pedestrians and vehicles on the crosswalk. At the same time, an emotion engine analyzes the user's emotional state and highlights information that is considered particularly dangerous if the user feels surprised. A projection device projects this information onto the user's retina, allowing the user to safely understand the situation at the crosswalk.

[1463] Example prompts to input to the generative AI model

[1464] "Please tell me a program that can analyze video data in real time on an AI analysis server and recognize specific objects."

[1465] "How can I use a sentiment analysis engine to detect a user's emotional state in real time?"

[1466] "Please tell me the specific steps for correcting image data with smart glasses and using a retinal projection device."

[1467] In this way, the invention can provide visually impaired users with optimal visual information according to their emotions, and can support safe and comfortable travel in cooperation with autonomous vehicles.

[1468] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1469] System processing steps

[1470] Step 1: Capture footage

[1471] The device uses a camera to capture visual information in real time. Specifically, a camera mounted on smart glasses worn by the user captures images of the scene ahead and stores them as digital data frame by frame.

[1472] Input: Front view of the user

[1473] Output: Captured video frame data

[1474] Step 2: Sending video data

[1475] The device transmits the captured video frame data to an AI analysis server via the Internet, enabling remote, high-performance analysis.

[1476] Input: Video frame data

[1477] Output: Send data to the AI ​​analysis server

[1478] Step 3: Analyzing the video data

[1479] The server processes the received video frame data in real time, using an AI model (e.g., the YOLO object recognition model, analysis using TensorFlow) to recognize objects and characters in the video, identify objects such as vehicles and traffic lights, and assign labels and bounding boxes to them.

[1480] Input: Transmitted video frame data

[1481] Output: Object and character recognition results (labels, bounding boxes)

[1482] Step 4: Recognizing your emotional state

[1483] The server uses an emotion engine to analyze the user's emotional state in real time, using voice analysis and facial expression recognition technology, e.g., Microsoft Azure Emotion API.

[1484] Input: User's voice and facial expression data

[1485] Output: Emotional state recognition result (surprise, anxiety, joy, etc.)

[1486] Step 5: Correcting the video data

[1487] The server generates an image that corrects for the visual field defect based on the analysis results, and processes the data by enlarging parts of the image to make important information easier for the user to see.

[1488] Input: Object recognition and character recognition results, emotional state recognition results

[1489] Output: Corrected video data

[1490] Step 6: Performing Highlighting

[1491] The server optimizes the video data correction and display content based on the user's emotional state detected by the emotion engine. For example, if the user is feeling anxious, important warning information will be highlighted.

[1492] Input: Corrected video data, emotional state recognition results

[1493] Output: Video data optimized for the user's emotional state

[1494] Step 7: Transmitting and projecting video data

[1495] The server then transmits the optimized image data based on the correction and emotion information to the terminal, which then projects the received corrected image data onto the user's retina in real time, providing the user with the necessary visual information.

[1496] Input: Optimized video data

[1497] Output: Visual information projected onto the retina

[1498] Example processing steps

[1499] Example 1: Traffic light recognition and highlighting

[1500] Step 1: The device captures the image of the traffic light ahead.

[1501] Input: Traffic light image

[1502] Output: Captured video frame data

[1503] Step 2: The device sends the captured video data to the AI ​​analysis server.

[1504] Input: Video frame data

[1505] Output: Send data to the AI ​​analysis server

[1506] Step 3: The server recognizes the traffic light as an object and identifies its color (red, yellow, green) and location.

[1507] Input: Transmitted video frame data

[1508] Output: Object recognition and traffic light color / position data

[1509] Step 4: The server recognizes the user's anxiety state using the emotion engine.

[1510] Input: User's voice and facial expression data

[1511] Output: Emotional state recognition result (anxiety)

[1512] Step 5: The server generates a corrected image by highlighting the traffic light information.

[1513] Input: Traffic light color / position data, emotional state recognition results

[1514] Output: Highlighted corrected video data

[1515] Step 6: The server sends the highlighted corrected video data to the terminal.

[1516] Input: Highlighted corrected video data

[1517] Output: Sending data to the terminal

[1518] Step 7: The device projects the highlighted corrected image data onto the user's retina.

[1519] Input: Highlighted corrected video data

[1520] Output: Visual information projected onto the retina

[1521] Example prompts to input to the generative AI model

[1522] "Please tell me a program that can analyze video data in real time on an AI analysis server and recognize specific objects."

[1523] "How can I use a sentiment analysis engine to detect a user's emotional state in real time?"

[1524] "Please tell me the specific steps for correcting image data with smart glasses and using a retinal projection device."

[1525] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1526] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1527] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1528] [Fourth embodiment]

[1529] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1530] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1531] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1532] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1533] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1534] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1535] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1536] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1537] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1538] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1539] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1540] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1541] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1542] An embodiment of this invention provides a system that complements visual information for users with impaired vision due to glaucoma, etc. This system is centered around a glasses-type device worn by the user, and is composed of a camera, an AI analysis server, and a retinal projection device.

[1543] System configuration

[1544] 1. Camera:

[1545] It is installed in glasses worn by the user and captures images in front of the vehicle in real time, which are then saved as image frame data.

[1546] 2. AI analysis server:

[1547] The system receives video data sent from the camera and analyzes it using an AI model. Specifically, it recognizes objects and characters in the video and identifies the information.

[1548] To compensate for the visual field defects, the analysis results are used to generate corrected video data, which includes scaling the video and adding text information.

[1549] 3. Retinal projection device:

[1550] This device projects corrected image data onto the user's retina in real time, allowing the user to compensate for visual field defects and easily obtain visual information in everyday life.

[1551] Program processing (natural language explanation)

[1552] Video capture and transmission

[1553] Device (camera): The camera in the glasses worn by the user captures the image in front of the user in real time. This image is saved as data frame by frame.

[1554] Terminal: Sends captured video data to an AI analysis server via the internet.

[1555] Video data analysis

[1556] Server: The AI ​​analytics server processes the received video data in real time, using AI models to recognize objects in the video, such as cars, park benches, and traffic lights, and assigns bounding boxes to each.

[1557] Server: Using OCR technology, character information in the video is extracted, obtaining text data such as "Watch out for pedestrians" or "The traffic light is red."

[1558] Video data correction

[1559] Server: Based on the analysis results, an image is generated to compensate for the visual field defect. Important information is reconstructed in a form that is easy for the user to recognize, using techniques such as image enlargement.

[1560] Server: Object information and extracted text information are overlaid on the video in a simple format.

[1561] Video data transmission and projection

[1562] Server: Sends the corrected video data to the user terminal.

[1563] Terminal: The transmitted corrected image data is projected onto the user's retina in real time. This process complements the missing information in the user's field of vision.

[1564] Specific examples

[1565] 1. If the user is waiting at a traffic light:

[1566] Terminal (camera): Captures the traffic light in front of the user.

[1567] Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, extracts the traffic light's text information (e.g., "pedestrian signal").

[1568] Server: Reconstructs the traffic light color and text information as a corrected image.

[1569] Device: Projects traffic light colors and text information onto the user's retina to complement visual information.

[1570] 2. If the user is walking down the street:

[1571] Terminal (camera): Captures images of the road and intersections ahead.

[1572] Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[1573] Server: Emphasizes information about pedestrians and vehicles that require attention and generates corrected images.

[1574] Terminal: This information is projected onto the user's retina to complement the visual information.

[1575] In this way, the system effectively supports the vision of glaucoma patients, helping them live safe and comfortable lives.

[1576] The processing flow will be explained below.

[1577] Step 1:

[1578] Device (camera): A camera mounted on glasses worn by the user captures images of the surroundings in real time, and these images are stored as digital data frame by frame.

[1579] Step 2:

[1580] Terminal: Captured video data is sent to the AI ​​analysis server via the internet using a low-latency, highly reliable communication protocol.

[1581] Step 3:

[1582] Server: Stores the received video data in a buffer for analysis.

[1583] Step 4:

[1584] Server: The AI ​​model analyzes the video data stored in the buffer and recognizes objects in the image, such as cars, pedestrians, and traffic lights, and assigns labels and bounding boxes to each.

[1585] Step 5:

[1586] Server: Analyzes text information contained in video using OCR technology and saves it as metadata. For example, extracts text from traffic lights and signs.

[1587] Step 6:

[1588] Server: Based on the analysis results, an image is generated that corrects for the visual field defect. Here, processing such as enlarging parts of the image is performed to make important information easier for the user to see.

[1589] Step 7:

[1590] Server: Object information and extracted text information are overlaid on the original video. For example, a large "STOP" message is displayed over a "red light."

[1591] Step 8:

[1592] Server: Converts the corrected video data into a format suitable for head-up displays and sends it to the user's device.

[1593] Step 9:

[1594] Terminal: The received corrected image data is projected onto the user's retina in real time, allowing the user to fill in information about the missing area of ​​their visual field.

[1595] Step 10:

[1596] User: Views the projected image and recognizes the surrounding situation and objects, enabling safe walking and driving.

[1597] Through this series of steps, the visual aid system effectively compensates for the user's visual field defects and provides the visual information necessary for daily life.

[1598] Example 1

[1599] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1600] Visually impaired users face challenges in moving safely and accurately grasping information about their surroundings in daily life. In particular, if they have lost vision due to diseases such as glaucoma, they are more likely to miss important visual information in the lost area, which could lead to serious accidents or danger.

[1601] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1602] In this invention, the server includes means for recognizing objects in the video, means for extracting text information, means for correcting the visual field defect and reconstructing the video, and means for enlarging the corrected video and adding text information, thereby enabling visually impaired users to supplement important information present in the visual field defect in real time and recognize their surroundings safely and reliably.

[1603] A "camera" is a photographic device for capturing images.

[1604] "Video data" refers to visual information captured by a camera and recorded in digital format.

[1605] "Artificial intelligence" is a technology in which computer programs mimic human intelligence to solve problems, and in this case it is used to analyze video data.

[1606] A "server" is a computer device that receives data over a network and analyzes and processes it.

[1607] A "projection device" is a device for projecting the analyzed and corrected image data onto the user's retina.

[1608] A "network" is an infrastructure for data communication, and is used here to transmit data between the camera and the server.

[1609] "Object recognition" is a technology that detects specific objects in video data and identifies their location and type.

[1610] "Text information" refers to characters and text data contained in video data.

[1611] "OCR" stands for Optical Character Recognition, a technology that extracts text information from video data.

[1612] The "visual field defect area" refers to an area that is missing from the user's field of vision due to a visual impairment.

[1613] "Image correction" is a correction process carried out to compensate for missing parts of the field of view based on analyzed image data.

[1614] "Real-time" refers to processing and display occurring immediately, without delay.

[1615] "Image enlargement" is a technique for displaying a specific image portion in a larger, more easily visible manner.

[1616] "Adding text information" is a technique for overlaying analyzed text information onto video.

[1617] This invention provides a system that complements visual information for users with visual impairments due to glaucoma, etc. The system consists of a camera, an AI analysis server, a network, and a projection device.

[1618] 1. Camera: The camera is mounted on a glasses-type device worn by the user and captures the image in front of the user in real time. The captured image is recorded as digital data frame by frame.

[1619] 2. AI analysis server: The video data captured by the camera is sent to the AI ​​analysis server via the network. The server performs the following processing steps:

[1620] Object Recognition: Using AI models to identify objects in a video (cars, traffic lights, pedestrians, etc.) and assign them bounding boxes, for example, identifying a traffic light as red.

[1621] Text information extraction: Extract text information (e.g., "pedestrian signal") from the video using OCR technology.

[1622] Image correction: To generate a corrected image, areas of vision defects are filled in and important information is emphasized, such as overlaying traffic light colors or important messages on the image.

[1623] 3. Network: Wi-Fi or 4G / 5G networks are used for data communication, sending and receiving data between the camera and the server.

[1624] 4. Projection device: Projects corrected image data onto the user's retina in real time, allowing the user to instantly see important information present in the area of ​​visual field loss.

[1625] Specific examples

[1626] 1. If the user is waiting at a traffic light:

[1627] Terminal (camera): Captures the traffic light in front of the user.

[1628] Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, extracts the traffic light's text information (e.g., "pedestrian signal").

[1629] Server: Reconstructs the traffic light color and text information as a corrected image.

[1630] Device: Projects traffic light colors and text information onto the user's retina to complement visual information.

[1631] 2. If the user is walking down the street:

[1632] Terminal (camera): Captures images of the road and intersections ahead.

[1633] Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[1634] Server: Emphasizes information about pedestrians and vehicles that require attention and generates corrected images.

[1635] Terminal: This information is projected onto the user's retina to complement the visual information.

[1636] Prompt Sentence Examples

[1637] Traffic light discrimination prompt: The prompt for extracting color and text information from traffic light images, generating corrected images for retinal projection is as follows: "Please distinguish between red, yellow, and green signals from traffic light images, extract the traffic light text information, and generate corrected images. Then, send the generated images to the retinal projection device."

[1638] Object Recognition Prompt: The prompt for recognizing cars, pedestrians, and traffic lights at intersections from the forward video while the user is walking and generating corrected video for safety information is as follows: "Please recognize cars, pedestrians, and traffic lights from the forward video while the user is walking, and reconstruct the corrected video for safety information. Then, send the generated video to the retinal projection device."

[1639] This system will enable glaucoma patients and other visually impaired users to effectively supplement their visual information in their daily lives, enabling them to live safe and comfortable lives.

[1640] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1641] Step 1:

[1642] The user turns on the eyeglasses-type device.

[1643] Input: Power-on operation

[1644] Specific operation: The user presses a button to turn on the device, which starts up the device's system and initializes the camera and retinal projection device.

[1645] Output: Device booted, camera and retinal projection device ready

[1646] Step 2:

[1647] The device (camera) captures the image in front of the user.

[1648] Input: Device boot complete, user visible

[1649] How it works: The camera is installed inside the device and captures real-time images of what is in front of the user's field of vision.

[1650] Output: Captured video data (frame by frame)

[1651] Step 3:

[1652] The device (camera) sends the captured video data to the server.

[1653] Input: Captured video data (frame by frame)

[1654] Specific operation: The captured video data is sent to an AI analysis server via Wi-Fi or 4G / 5G networks.

[1655] Output: Video data received by the server

[1656] Step 4:

[1657] The server inputs the received video data into the AI ​​model and begins analysis.

[1658] Input: Video data received by the server

[1659] How it works: The server inputs the video data into the AI ​​model and performs object recognition and OCR processing, such as identifying traffic lights, pedestrians, and text information and enclosing them in bounding boxes.

[1660] Output: Analyzed video data (object recognition information, text information)

[1661] Step 5:

[1662] The server generates an image that corrects the visual field defect based on the analysis results.

[1663] Input: Analyzed video data (object recognition information, text information)

[1664] How it works: The server uses the analysis results to generate an image that complements the visual field defect, enlarging important information and adding warning text such as "Watch your direction."

[1665] Output: Corrected video data

[1666] Step 6:

[1667] The server transmits the corrected video data to the user terminal.

[1668] Input: Corrected video data

[1669] Specific operation: The corrected video data is immediately sent to the user terminal and adjusted to minimize delay.

[1670] Output: Corrected video data received by the user device

[1671] Step 7:

[1672] The terminal (retinal projection device) projects the corrected image data onto the user's retina.

[1673] Input: Corrected video data received by the user terminal

[1674] How it works: The retinal projection device accurately projects the corrected image data into the user's field of vision, for example, displaying the red light of a traffic light and the text "Please stop" in the center of the field of vision.

[1675] Output: Corrected image projected onto the user's retina

[1676] In this way, each step of the system is performed sequentially, allowing visually impaired users to complete important visual information in real time.

[1677] (Application example 1)

[1678] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1679] This invention relates to a visual aid system that enables visually impaired users to appropriately acquire visual information and act safely in daily life and in specific situations. In particular, the objective is to provide a system that analyzes visual information in real time while driving a car and complements and emphasizes important information, thereby enabling visually impaired people to drive safely. Current visual aid systems lack analytical accuracy, real-time performance, and retinal projection technology, making improving driving safety a challenge.

[1680] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1681] In this invention, the server includes a means for capturing video using a camera, a computer equipped with artificial intelligence for analyzing the captured video, a projection device for correcting the analyzed video data and projecting it onto the user's retina, a means for transmitting the corrected video data from the computer to a terminal in real time, and a means for displaying the corrected video data on the terminal. This allows visually impaired users to accurately grasp visual information while driving, significantly improving safety. Furthermore, real-time analysis and visual enhancement allow users to immediately recognize important information such as obstacles and traffic signals, shortening the user's reaction time and making it easier to avoid danger.

[1682] A "camera" is a device for capturing images and is a piece of equipment used to record visual information.

[1683] "Artificial intelligence for video analysis" is a computer system that processes captured video data and performs object and character recognition.

[1684] A "computer" is a device that processes and analyzes data, and in particular, a computer equipped with artificial intelligence (AI)-based image analysis technology.

[1685] A "projection device" is a device that has the technology to project analyzed image data onto the user's vision, and in particular refers to equipment that projects images directly onto the retina.

[1686] "Real-time transmission means" refers to the technology and protocols for immediately transmitting processed data to the terminal.

[1687] A "terminal" is a device for displaying received video data, and includes smart glasses, head-mounted displays, and the like.

[1688] "Means for recognizing objects" refers to technology for detecting objects and people contained in video and identifying their location and type.

[1689] "Means for extracting text information" refers to OCR (optical character recognition) technology for analyzing and reading text data contained in video.

[1690] "Means for reconstructing images" refers to technology for correcting visual field defects and regenerating image data in a form that is easy for users to understand.

[1691] A "visual display device" is a device for visually presenting analyzed and corrected image data to a user, including smart glasses and retinal projection devices.

[1692] As an embodiment of the present invention, a system is provided for supplementing visual information for users with visual impairments such as glaucoma while driving a car. This system includes a camera, a computer (AI analysis server), a projection device, and a terminal. The role of each component and the overall processing flow are described in detail below.

[1693] System configuration

[1694] 1. Camera

[1695] The camera captures real-time images of the driving environment, which are then stored as data frame by frame and sent to a computer. The camera is used in smart glasses or head-mounted displays.

[1696] 2. AI analysis server

[1697] The computer (AI analysis server) receives the video data sent from the camera and analyzes it using an artificial intelligence model. Specifically, it recognizes objects in the video (cars, pedestrians, traffic lights, etc.) and extracts text information (intersection signs, road information, etc.). It then corrects for any visual defects and reconstructs the video in a way that emphasizes important information.

[1698] 3. Projection device

[1699] The corrected image data is projected directly onto the user's retina through a projection device, such as a device inside a pair of smart glasses, to complement the user's visual information in real time.

[1700] 4. Terminal

[1701] The corrected video data is sent in real time from the AI ​​analysis server to the device that displays the received video data, such as smart glasses or a head-mounted display.

[1702] Operational Overview

[1703] The specific steps involved in the operation of this system are described below.

[1704] 1. Video capture and transmission

[1705] The camera captures the image in front of the user and saves it as video frame data, which is then sent to an AI analysis server via the internet.

[1706] Example prompt sentence:

[1707] "Capture the video frame data to send to the AI ​​server and send it over the Internet. The video frame format is JPEG, and the server URL is 'http: / / ai-server / parse_frame'."

[1708] 2. Analysis and correction of video data

[1709] The server processes the received video data in real time. Using AI analysis, it recognizes objects in the video, identifying things like cars, park benches, and traffic lights. It also uses optical character recognition (OCR) technology to extract text information from the video. Based on the analysis results, it generates an image to complement any visual field defects, visually emphasizing important information.

[1710] Example prompt sentence:

[1711] "Please recognize objects and text contained in the received video data and return that information as a result. We use the YOLO model for object recognition and OCR technology for character recognition."

[1712] 3. Projecting the corrected image

[1713] The corrected image data sent from the server is received by the device and projected onto the user's retina in real time, allowing the user to supplement their visual information and accurately recognize important information while driving.

[1714] Example prompt sentence:

[1715] "Send the corrected image data to the retinal projection device so that the user can accurately perceive the visual information."

[1716] Specific examples

[1717] 1. Waiting at a traffic light

[1718] When the user approaches an intersection, the camera captures the traffic lights and intersection signs, and the AI ​​analysis server recognizes the color of the traffic lights (red, yellow, green) and provides that information to the user.

[1719] Example prompt sentence:

[1720] "Analyze the color and text information of traffic lights, and if the light is yellow, display 'Proceed with caution'."

[1721] 2. Walking on the street

[1722] As a user crosses the road, the camera captures the movements of pedestrians and vehicles. The AI ​​analytics server analyzes this information, highlights areas that require attention, and provides the user with enhanced video.

[1723] Example prompt sentence:

[1724] "Analyze the location of pedestrians and other vehicles as you approach an intersection and highlight any potentially dangerous situations."

[1725] This system will enable visually impaired people to obtain the visual information they need safely while driving a car, significantly improving driving safety.

[1726] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1727] Step 1:

[1728] A camera captures an image in front of the user.

[1729] Input: The user's current field of view.

[1730] Output: Video frame data.

[1731] How it works: The camera built into the smart glasses or head-mounted display captures video in real time and generates video frame data.

[1732] Step 2:

[1733] The device sends the captured video data to an AI analysis server.

[1734] Input: Video frame data.

[1735] Output: Video data sent to the server.

[1736] Specific operation: The device sends video data via the Internet, and the prompt to be entered is "Capture the video frame data to be sent to the AI ​​server and send it via the Internet. The video frame format is JPEG, and the server URL is 'http: / / ai-server / parse_frame'."

[1737] Step 3:

[1738] The server analyzes the received video data.

[1739] Input: Transmitted video data.

[1740] Output: Analysis results (object and character recognition data).

[1741] Specific operation: The server uses a generative AI model such as the YOLO model to recognize objects and text contained in the video data and identify information such as traffic lights, pedestrians, and signs. An example of a prompt is "Please recognize the objects and text contained in the received video data and return that information as a result. The YOLO model is used for object recognition, and OCR technology is used for character recognition."

[1742] Step 4:

[1743] The server corrects the video based on the analysis results.

[1744] Input: Analysis results.

[1745] Output: Corrected video data.

[1746] How it works: Based on the analysis results, the server compensates for the visual field defects and reconstructs the image to highlight important information. This correction makes the color of traffic lights and the content of signs appear larger, for example.

[1747] Step 5:

[1748] The server transmits the corrected video data to the terminal in real time.

[1749] Input: Corrected video data.

[1750] Output: Corrected video data sent to the device.

[1751] Specific operation: The server sends the corrected image data to the terminal in real time, and an example of a prompt is used: "Please send the corrected image data to the retinal projection device so that the user can accurately perceive the visual information."

[1752] Step 6:

[1753] The terminal displays the corrected image data on a display device (retinal projection device).

[1754] Input: Corrected video data.

[1755] Output: The corrected image displayed in the user's field of view.

[1756] How it works: The device sends the corrected image data to the retinal projection device, which projects it into the user's field of vision. For example, traffic light colors and sign information are projected directly onto the user's retina, allowing the user to perceive the necessary information, even if they have visual impairments.

[1757] By following the steps above, a visually impaired user can accurately obtain visual information while driving a car, enabling safe driving.

[1758] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1759] An embodiment of this invention provides a system that complements visual information for users with impaired vision due to glaucoma, etc., and also provides optimal auxiliary information according to the user's emotional state. This system is centered around a glasses-type device worn by the user, and is composed of a camera, an AI analysis server, an emotion engine, and a retinal projection device.

[1760] System configuration

[1761] 1. Camera:

[1762] The device is installed in glasses worn by the user and captures images of the surrounding area in real time, which are then stored as digital data frame by frame.

[1763] 2. AI analysis server:

[1764] The system receives video data sent from the camera and analyzes it using an AI model. Specifically, it recognizes objects and characters in the video and identifies the information.

[1765] To compensate for the visual field defects, the analysis results are used to generate corrected video data, which includes scaling the video and adding text information.

[1766] 3. Emotion Engine:

[1767] Recognize the user's emotional state in real time. The emotion engine can use technologies such as facial expression recognition and voice analysis.

[1768] Adjusting the display of visual information depending on emotional state (e.g., surprise, anxiety, joy, etc.).

[1769] 4. Retinal projection device:

[1770] This device projects corrected image data onto the user's retina in real time, allowing the user to compensate for visual field defects and easily obtain visual information in everyday life.

[1771] Program processing (natural language explanation)

[1772] Video capture and transmission

[1773] Device (camera): The camera in the glasses worn by the user captures the image in front of the user in real time, and this image is saved as digital data frame by frame.

[1774] Terminal: Sends captured video data to an AI analysis server via the internet.

[1775] Video data analysis

[1776] Server: The AI ​​analytics server processes the received video data in real time, using AI models to recognize objects in the video, such as cars, park benches, and traffic lights, and assigns labels and bounding boxes to each.

[1777] Server: Using OCR technology, the text information in the video is analyzed and text data such as "Watch out for pedestrians" or "The traffic light is red" is obtained.

[1778] Video data correction

[1779] Server: Based on the analysis results, an image is generated that corrects for the visual field defect. Here, processing such as enlarging parts of the image is performed to make important information easier for the user to see.

[1780] Server: Based on the object information and extracted text information, the server reconstructs the image with appropriate correction processing.

[1781] Emotion engine processing

[1782] Server: Analyzes the user's emotional state in real time using an emotion engine, for example, by using technology to detect emotions from facial expressions and voice.

[1783] Server: Based on the user's emotional state detected by the emotion engine, the server corrects the video data and optimizes the display content. For example, if the user is feeling anxious, it can highlight warning information.

[1784] Video data transmission and projection

[1785] Server: Transmits optimized video data based on correction and emotion information to the terminal.

[1786] Terminal: The received corrected image data is projected onto the user's retina in real time, allowing the user to compensate for the visual field defect and obtain appropriate visual information corresponding to their emotions.

[1787] Specific examples

[1788] 1. If the user is waiting at a traffic light:

[1789] Terminal (camera): Captures the traffic light in front of the user.

[1790] Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, analyzes the text displayed on the traffic lights.

[1791] Server: Generates corrected images based on the color and text information of traffic lights.

[1792] Server: If the emotion engine recognizes the user's state of anxiety, it highlights the traffic light information.

[1793] Device: Projects corrected and highlighted traffic light information onto the user's retina, allowing the user to see the traffic light with confidence.

[1794] 2. If the user is walking down the street:

[1795] Terminal (camera): Captures images of the road and intersections ahead.

[1796] Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[1797] Server: Based on the analysis results, the server generates corrected images that emphasize information that requires attention.

[1798] Server: If the emotion engine recognizes the user's surprise or fear, it prioritizes displaying information that is considered particularly dangerous.

[1799] Terminal: This information is projected onto the retina, allowing the user to respond appropriately.

[1800] In this way, the system effectively supports the vision of glaucoma patients and also supports a safe and comfortable life by providing appropriate information according to their emotions.

[1801] The processing flow will be explained below.

[1802] Step 1:

[1803] Device (camera): A camera mounted on glasses worn by the user captures images of the surroundings in real time, and these images are stored as digital data frame by frame.

[1804] Step 2:

[1805] Terminal: Captured video data is sent to the AI ​​analysis server via wireless communication, using a low-latency, highly reliable communication protocol.

[1806] Step 3:

[1807] Server: Stores the received video data in a buffer for analysis.

[1808] Step 4:

[1809] Server: The AI ​​model analyzes the video data stored in the buffer and recognizes objects in the image, such as cars, pedestrians, and traffic lights, and assigns labels and bounding boxes to each.

[1810] Step 5:

[1811] Server: Analyzes textual information contained in video using OCR technology and saves it as metadata. For example, extracts text from traffic lights and signs.

[1812] Step 6:

[1813] Server: Analyzes the user's emotional state in real time using an emotion engine. It uses facial expression recognition and voice analysis technology to detect emotional states (e.g., surprise, anxiety, joy, etc.).

[1814] Step 7:

[1815] Server: Based on the analysis results and the user's emotional state, an image is generated that corrects for the visual field defect. This involves enlarging parts of the image to make important information easier for the user to see.

[1816] Step 8:

[1817] Server: Further adjusts the correction content of the video data based on the user's emotional state recognized by the emotion engine. For example, if the user is feeling anxious, highlight warning information.

[1818] Step 9:

[1819] Server: Object information and extracted text information are overlaid on the original video. For example, a large "STOP" message is displayed over a "red light."

[1820] Step 10:

[1821] Server: Transmits video data optimized based on correction and emotion information to the terminal.

[1822] Step 11:

[1823] Terminal: The received corrected image data is projected onto the user's retina in real time, allowing the user to fill in information about the missing area of ​​their visual field.

[1824] Step 12:

[1825] User: Views the projected image and recognizes the surrounding situation and objects, enabling safe walking and driving.

[1826] Specific examples

[1827] When the user is waiting at a traffic light

[1828] 1. Device (camera): Captures the traffic light in front of the user.

[1829] 2. Terminal: Sends the captured video data to the server.

[1830] 3. Server: Recognizes traffic lights as objects and identifies their colors (red, yellow, green). At the same time, analyzes the text displayed on the traffic lights.

[1831] 4. Server: Generates corrected images based on the traffic light color and text information.

[1832] 5. Server: If the emotion engine recognizes the user's anxious state, it highlights the traffic light information.

[1833] 6. Server: Sends the corrected and highlighted traffic light information to the terminal.

[1834] 7. Terminal: The corrected image is projected onto the user's retina, allowing the user to see the signal with confidence.

[1835] If the user is walking down the street

[1836] 1. Terminal (camera): Captures images of the road and intersection ahead.

[1837] 2. Terminal: Sends video data to the server.

[1838] 3. Server: Recognizes pedestrians, cars, traffic lights, etc. and identifies their locations and movements in real time.

[1839] 4. Server: Based on the analysis results, the server generates a corrected image by emphasizing information that requires attention.

[1840] 5. Server: If the emotion engine recognizes the user's surprise or fear, it prioritizes displaying information that is considered particularly dangerous.

[1841] 6. Server: Sends the corrected and enhanced information to the terminal.

[1842] 7. Terminal: This information is projected onto the retina, allowing the user to respond appropriately.

[1843] Through this series of steps, the visual assistance system effectively compensates for the user's visual field loss and also provides appropriate information according to their emotions, supporting a safe and comfortable life.

[1844] Example 2

[1845] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1846] Conventional visual assistance systems compensate for visual field defects in visually impaired people without taking the user's emotional state into consideration, which can result in information overload or overlooking important information. Furthermore, the provision of information adapted to the user's behavioral situation and surrounding environment is insufficient, making it difficult for the user to act safely and securely. This has led to a need for improved safety and comfort in everyday life.

[1847] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1848] In this invention, the server includes means for analyzing video data, means for correcting visual field defects, and means for analyzing the user's emotional state, thereby enabling the correction of video data and optimization of display content based on the user's emotional state.

[1849] A "camera" is a device for capturing images, and is built into a glasses-type device worn by a user.

[1850] An "AI-equipped server" is a computer system that receives and analyzes captured video data, and is a device equipped with an AI model that performs object and character recognition.

[1851] A "projection device" is a device for projecting corrected image data onto the user's retina in real time.

[1852] "Means for analyzing emotional state" refers to technology or devices for analyzing the user's facial expressions, voice, etc. to determine the user's emotional state in real time.

[1853] "Means for recognizing objects" refers to technology that identifies objects in video and assigns labels and bounding boxes to them.

[1854] "Means for extracting text information" refers to technology that analyzes text in video and obtains text data.

[1855] "Means to correct visual field defects" refers to technologies that enlarge or reduce images or overlay text information to compensate for the user's visual field defects.

[1856] "Means for generating video data in real time" refers to technology for instantly generating corrected video based on analyzed data.

[1857] The embodiment of this invention is a system that complements visual information for visually impaired users and provides optimal auxiliary information according to the user's emotional state. This system is composed of a glasses-type device at its core, a camera, an AI analysis server, an emotion engine, and a retinal projection device.

[1858] System Configuration

[1859] 1. Camera:

[1860] The camera is embedded in glasses worn by the user and captures images of the area in front of it in real time, capturing images at 30 frames per second and storing them as digital data.

[1861] 2. AI analysis server:

[1862] The server receives video data sent from the device and analyzes it using an AI model. Specifically, it uses deep learning models such as ResNet and YOLO to recognize objects in the video, assign labels and bounding boxes, and use OCR technology to extract text information from the video.

[1863] 3. Emotion Engine:

[1864] The server uses an emotion engine to analyze the user's emotional state in real time, which uses technology to analyze facial expressions and voice to determine the user's emotional state (surprise, anxiety, joy, etc.).

[1865] 4. Retinal projection device:

[1866] The device projects the corrected image data onto the user's retina in real time, allowing the user to compensate for the visual field defect and obtain optimal visual information according to their emotions.

[1867] Program processing details

[1868] Video capture and transmission:

[1869] When a user puts on the glasses, the camera activates and starts capturing images of the surrounding area, which are then sent to an AI analysis server via Wi-Fi, 4G, or 5G networks.

[1870] Video data analysis and correction:

[1871] The server analyzes the received video data, performs object and character recognition, and then makes corrections based on the analysis results, such as enlarging a portion of the video.

[1872] Emotional State Analysis:

[1873] The server uses an emotion engine to determine the user's emotional state and highlights information according to anxiety, surprise, etc.

[1874] Generate and project corrected image data:

[1875] The server generates the corrected image data and transmits it to the terminal, which then projects it into the user's field of vision via a retinal projection device.

[1876] Specific example of operation

[1877] 1. If the user is waiting at a traffic light:

[1878] The device (camera) captures images of traffic lights and sends them to the server.

[1879] The server analyzes the color and text displayed on the traffic light and generates a corrected image.

[1880] The server adjusts the display to highlight traffic light information because the emotion engine recognizes the user's anxiety.

[1881] The device projects highlighted traffic light information onto the user's retina, allowing them to see the traffic lights with confidence.

[1882] 2. If the user is walking down the street:

[1883] The device (camera) captures images of the road and intersection ahead and sends them to the server.

[1884] The server recognizes pedestrians, cars and traffic lights and determines their location and movement.

[1885] The server generates corrected images that highlight important information, and if the emotion engine recognizes surprise or fear, it prioritizes displaying dangerous information.

[1886] The device projects highlighted information onto the user's retina, helping them navigate safely.

[1887] Examples of prompts for generative AI models:

[1888] "If a user with impaired vision due to glaucoma or other conditions is waiting at a traffic light, the emotion engine should generate optimally corrected video data when it recognizes the user's state of anxiety."

[1889] In this way, a system is realized that provides optimal visual information to users and supports safe and comfortable daily life.

[1890] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1891] Step 1:

[1892] Device: The user wears the glasses-type device. The built-in camera automatically starts up and captures the image in front of the user in real time at a rate of 30 frames per second. The captured image is stored in the internal memory as digital data.

[1893] Input: Camera video (analog signal)

[1894] Output: Digital video data (real time)

[1895] Step 2:

[1896] Terminal: Captured video data is packaged frame by frame and sent to an AI analysis server via Wi-Fi, 4G, or 5G networks. Data is sent using a low-latency communication protocol and adaptive data transfer technology according to the transmission conditions.

[1897] Input: Digital video data (internal memory)

[1898] Output: Data packets sent over the network

[1899] Step 3:

[1900] Server: Receives the transmitted video data and loads it into a data queue. Then, it uses a deep learning model (e.g., ResNet, YOLO) to analyze objects in the video in real time. During the analysis process, objects are assigned labels and bounding boxes.

[1901] Input: Video data received via the network

[1902] Output: Object recognition data (labels, bounding boxes)

[1903] Step 4:

[1904] Server: Analyzes text information in video using OCR technology. Specifically, it extracts text data from traffic lights, road signs, etc.

[1905] Input: Analyzed video data

[1906] Output: Text data (e.g., "Watch out for pedestrians," "The traffic light is red")

[1907] Step 5:

[1908] Server: Based on the analysis results, the image of the visual field defect is corrected. Part of the image is enlarged to make important information easier to see, and necessary text information is overlaid. During this process, warnings and cautions are highlighted.

[1909] Input: object recognition data, text data

[1910] Output: Corrected video data

[1911] Step 6:

[1912] Server: Analyzes the user's emotional state using an emotion engine. It uses voice and facial expression data acquired from the device's built-in microphone and camera to determine the user's emotions (e.g., surprise, anxiety, joy).

[1913] Input: User's voice and facial expression data

[1914] Output: Emotional state data

[1915] Step 7:

[1916] Server: Integrates emotional state data and video analysis data to generate optimally corrected video data. If the user feels anxious, it generates video that highlights warning information.

[1917] Input: Corrected video data, emotional state data

[1918] Output: Optimized corrected video data

[1919] Step 8:

[1920] Server: Sends the generated optimized corrected image data to the terminal. The data is compressed and transmitted quickly and efficiently.

[1921] Input: Optimized corrected image data

[1922] Output: Data packets sent over the network

[1923] Step 9:

[1924] Terminal: The received corrected image data is projected onto the user's retina in real time. The retinal projection device visually presents information to compensate for the visual field defect, allowing the user to act safely and effectively.

[1925] Input: Received corrected video data

[1926] Output: Visual information projected onto the retina

[1927] (Application example 2)

[1928] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1929] Existing visual assistance systems are limited to supplementing visual information for users with visual impairments and do not take into account the user's emotional state. Furthermore, they lack the ability to link with autonomous vehicles and provide real-time safety information inside and outside the vehicle. This has resulted in insufficient safety and comfort for visually impaired people in certain environments.

[1930] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1931] In this invention, the server includes a means for capturing video using a camera, a means equipped with AI for analyzing the captured video, a means for correcting the analyzed video data and projecting it onto the user's retina, a means for linking with an emotion engine that recognizes the user's emotional state and adjusts the information display, and a means for linking with an autonomous vehicle and presenting important information inside and outside the vehicle. This makes it possible to provide optimal visual information according to the user's emotional state, and by linking with an autonomous vehicle, it is possible to improve the safety and comfort of travel for visually impaired people.

[1932] A "camera" is a device for capturing visual information, acquiring video and images as digital data.

[1933] An "AI-equipped server" is a server that uses artificial intelligence technology to analyze video data and has the ability to recognize objects and characters.

[1934] A "projection device" is a device for projecting corrected image data directly onto the user's visual organs (such as the retina).

[1935] The "emotion engine" is an engine that recognizes the user's emotional state and adjusts the operation and display content of other systems based on that information.

[1936] An "autonomous vehicle" is a vehicle that can navigate roads autonomously without the need for human operation.

[1937] "Object recognition" is a technology that detects objects in an image and identifies what they are.

[1938] "Means for extracting text information" refers to technology that analyzes and extracts text from video.

[1939] "Correction of visual field defects" is a technology that allows visually impaired users to complement the parts they cannot see and reconstruct images.

[1940] "Means for highlighting information" refers to a technique for visually highlighting and displaying important information.

[1941] "Means of generating data in real time" refers to technology that instantly analyzes and processes acquired data and immediately displays or outputs the results.

[1942] "Data integration" refers to technology that allows data to be shared between different systems and utilized mutually.

[1943] To realize this application example, a visual aid system is required. The core system configuration is as follows:

[1944] System configuration

[1945] 1. Camera

[1946] Overview: Visual information is captured in real time using a camera mounted on a device worn by the user (such as smart glasses).

[1947] Hardware used: Camera in smart glasses (e.g. Google Glass)

[1948] 2. AI-powered servers

[1949] Overview: Analyzes video data sent from a camera and performs object and character recognition. Artificial intelligence technology is used for the specific analysis.

[1950] Software and hardware used: AI analysis server (e.g., AWS EC2, Google Cloud AI)

[1951] 3. Projection device

[1952] Overview: A device that projects corrected image data onto the user's retina, compensating for the user's visual impairment.

[1953] Hardware used: Retinal projection device in smart glasses

[1954] 4. Emotion Engine

[1955] Overview: Analyzes the user's emotional state in real time and displays the most appropriate information based on that emotion. Uses facial expression recognition and voice analysis technology.

[1956] Software used: Sentiment analysis engine (e.g. Microsoft Azure Emotion API)

[1957] 5. Collaboration with autonomous vehicles

[1958] Overview: Connecting data to autonomous vehicles provides users with important information about the inside and outside of the vehicle.

[1959] Hardware and software used: Communication interfaces for autonomous vehicles

[1960] Specific use cases

[1961] Example 1: Traffic light recognition and highlighting

[1962] When the user approaches an intersection, the camera captures the traffic light ahead. This video data is sent to an AI analysis server, which recognizes the color and position of the traffic light. At the same time, the emotion engine analyzes the user's emotional state and highlights the traffic light information if the user is feeling anxious. A projection device projects this information onto the user's retina, allowing the user to see the highlighted traffic light information.

[1963] Example 2: Autonomous vehicle crossing a pedestrian crossing

[1964] When a user is in an autonomous vehicle, a camera captures video of the crosswalk ahead. This video data is sent to an AI analysis server, which recognizes the movements of pedestrians and vehicles on the crosswalk. At the same time, an emotion engine analyzes the user's emotional state and highlights information that is considered particularly dangerous if the user feels surprised. A projection device projects this information onto the user's retina, allowing the user to safely understand the situation at the crosswalk.

[1965] Example prompts to input to the generative AI model

[1966] "Please tell me a program that can analyze video data in real time on an AI analysis server and recognize specific objects."

[1967] "How can I use a sentiment analysis engine to detect a user's emotional state in real time?"

[1968] "Please tell me the specific steps for correcting image data with smart glasses and using a retinal projection device."

[1969] In this way, the invention can provide visually impaired users with optimal visual information according to their emotions, and can support safe and comfortable travel in cooperation with autonomous vehicles.

[1970] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1971] System processing steps

[1972] Step 1: Capture footage

[1973] The device uses a camera to capture visual information in real time. Specifically, a camera mounted on smart glasses worn by the user captures images of the scene ahead and stores them as digital data frame by frame.

[1974] Input: Front view of the user

[1975] Output: Captured video frame data

[1976] Step 2: Sending video data

[1977] The device transmits the captured video frame data to an AI analysis server via the Internet, enabling remote, high-performance analysis.

[1978] Input: Video frame data

[1979] Output: Send data to the AI ​​analysis server

[1980] Step 3: Analyzing the video data

[1981] The server processes the received video frame data in real time, using an AI model (e.g., the YOLO object recognition model, analysis using TensorFlow) to recognize objects and characters in the video, identify objects such as vehicles and traffic lights, and assign labels and bounding boxes to them.

[1982] Input: Transmitted video frame data

[1983] Output: Object and character recognition results (labels, bounding boxes)

[1984] Step 4: Recognizing your emotional state

[1985] The server uses an emotion engine to analyze the user's emotional state in real time, using voice analysis and facial expression recognition technology, e.g., Microsoft Azure Emotion API.

[1986] Input: User's voice and facial expression data

[1987] Output: Emotional state recognition result (surprise, anxiety, joy, etc.)

[1988] Step 5: Correcting the video data

[1989] The server generates an image that corrects for the visual field defect based on the analysis results, and processes the data by enlarging parts of the image to make important information easier for the user to see.

[1990] Input: Object recognition and character recognition results, emotional state recognition results

[1991] Output: Corrected video data

[1992] Step 6: Performing Highlighting

[1993] The server optimizes the video data correction and display content based on the user's emotional state detected by the emotion engine. For example, if the user is feeling anxious, important warning information will be highlighted.

[1994] Input: Corrected video data, emotional state recognition results

[1995] Output: Video data optimized for the user's emotional state

[1996] Step 7: Transmitting and projecting video data

[1997] The server then transmits the optimized image data based on the correction and emotion information to the terminal, which then projects the received corrected image data onto the user's retina in real time, providing the user with the necessary visual information.

[1998] Input: Optimized video data

[1999] Output: Visual information projected onto the retina

[2000] Example processing steps

[2001] Example 1: Traffic light recognition and highlighting

[2002] Step 1: The device captures the image of the traffic light ahead.

[2003] Input: Traffic light image

[2004] Output: Captured video frame data

[2005] Step 2: The device sends the captured video data to the AI ​​analysis server.

[2006] Input: Video frame data

[2007] Output: Send data to the AI ​​analysis server

[2008] Step 3: The server recognizes the traffic light as an object and identifies its color (red, yellow, green) and location.

[2009] Input: Transmitted video frame data

[2010] Output: Object recognition and traffic light color / position data

[2011] Step 4: The server recognizes the user's anxiety state using the emotion engine.

[2012] Input: User's voice and facial expression data

[2013] Output: Emotional state recognition result (anxiety)

[2014] Step 5: The server generates a corrected image by highlighting the traffic light information.

[2015] Input: Traffic light color / position data, emotional state recognition results

[2016] Output: Highlighted corrected video data

[2017] Step 6: The server sends the highlighted corrected video data to the terminal.

[2018] Input: Highlighted corrected video data

[2019] Output: Sending data to the terminal

[2020] Step 7: The device projects the highlighted corrected image data onto the user's retina.

[2021] Input: Highlighted corrected video data

[2022] Output: Visual information projected onto the retina

[2023] Example prompts to input to the generative AI model

[2024] "Please tell me a program that can analyze video data in real time on an AI analysis server and recognize specific objects."

[2025] "How can I use a sentiment analysis engine to detect a user's emotional state in real time?"

[2026] "Please tell me the specific steps for correcting image data with smart glasses and using a retinal projection device."

[2027] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2028] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2029] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2030] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2031] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2032] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2033] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2034] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2035] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2036] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2037] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2038] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2039] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2040] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2041] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2042] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2043] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2044] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2045] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2046] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2047] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2048] The following is further disclosed regarding the above embodiment.

[2049] (Claim 1)

[2050] a means for capturing video by a camera;

[2051] A server equipped with AI that analyzes the captured video,

[2052] a projection device that corrects the analyzed image data and projects it onto the user's retina;

[2053] Visual aid systems including:

[2054] (Claim 2)

[2055] a means for recognizing objects in the video;

[2056] A means for extracting textual information;

[2057] A means for correcting a visual field defect and reconstructing an image is provided.

[2058] 10. The visual aid system of claim 1.

[2059] (Claim 3)

[2060] a means for generating image data to be projected onto a user's retina in real time;

[2061] 10. The visual aid system of claim 1.

[2062] "Example 1"

[2063] (Claim 1)

[2064] a means for capturing video by a camera;

[2065] A server equipped with artificial intelligence that analyzes the captured video,

[2066] a projection device that corrects the analyzed image data and projects it onto the user's retina;

[2067] means for transmitting the captured video data to a server via a network;

[2068] means for receiving the corrected video data from the server;

[2069] A system including:

[2070] (Claim 2)

[2071] a means for recognizing objects in the video;

[2072] A means for extracting textual information;

[2073] A means for correcting the visual field defect and reconstructing the image;

[2074] means for enlarging the corrected image and adding text information;

[2075] 10. The system of claim 1, comprising:

[2076] (Claim 3)

[2077] a means for generating image data to be projected onto a user's retina in real time;

[2078] 10. The system of claim 1.

[2079] "Application Example 1"

[2080] (Claim 1)

[2081] a means for capturing video by a camera;

[2082] A computer equipped with artificial intelligence to analyze the captured video,

[2083] a projection device that corrects the analyzed image data and projects it onto the user's retina;

[2084] means for transmitting the corrected video data from the computer to the terminal in real time;

[2085] means for displaying the corrected video data on the terminal;

[2086] A system including:

[2087] (Claim 2)

[2088] a means for recognizing objects in the video;

[2089] A means for extracting textual information;

[2090] A means for correcting a visual field defect and reconstructing an image is provided.

[2091] 10. The system of claim 1.

[2092] (Claim 3)

[2093] a means for generating image data in real time to be projected onto the user's retina;

[2094] A means for analyzing environmental information during driving and visually highlighting important information;

[2095] means for transmitting and displaying the corrected image data on a visual display device;

[2096] 10. The system of claim 1.

[2097] "Example 2: Combining Emotion Engines"

[2098] (Claim 1)

[2099] a means for capturing video by a camera;

[2100] A server means equipped with AI to analyze the captured video;

[2101] a projection device means for correcting the analyzed image data and projecting it onto the user's retina;

[2102] means for analyzing the emotional state of a user;

[2103] means for optimizing the correction and display of video data based on the emotional state;

[2104] A system including:

[2105] (Claim 2)

[2106] a means for recognizing objects in the video;

[2107] A means for extracting textual information;

[2108] A means for correcting a visual field defect and reconstructing an image is provided.

[2109] 10. The system of claim 1.

[2110] (Claim 3)

[2111] a means for generating image data to be projected onto a user's retina in real time;

[2112] 10. The system of claim 1.

[2113] "Application example 2 when combining emotion engines"

[2114] (Claim 1)

[2115] a means for capturing video by a camera;

[2116] A server equipped with AI that analyzes the captured video,

[2117] a projection device that corrects the analyzed image data and projects it onto the user's retina;

[2118] means for interfacing with an emotion engine to recognize an emotional state and adjust information display;

[2119] A means of interacting with autonomous vehicles and presenting important information inside and outside the vehicle;

[2120] Visual aid systems including:

[2121] (Claim 2)

[2122] a means for recognizing objects in the video;

[2123] A means for extracting textual information;

[2124] A means for correcting a visual field defect and reconstructing an image is provided.

[2125] means for highlighting information based on emotional state;

[2126] 10. The system of claim 1.

[2127] (Claim 3)

[2128] a means for generating image data to be projected onto a user's retina in real time;

[2129] Equipped with a means for data exchange with autonomous vehicles,

[2130] 10. The system of claim 1. [Explanation of symbols]

[2131] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for capturing video by a camera; A server equipped with AI that analyzes the captured video, a projection device that corrects the analyzed image data and projects it onto the user's retina; Visual aid systems including:

2. a means for recognizing objects in the video; A means for extracting textual information; A means for correcting a visual field defect and reconstructing an image is provided. The visual aid system of claim 1 .

3. a means for generating image data to be projected onto a user's retina in real time; The visual aid system of claim 1 .

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A