system
AR glasses with generative AI analysis and interactive feedback assist elderly users in operating complex devices, addressing interface complexity and enhancing digital device accessibility.
Patent Information
- Application Number
- JP2024138191
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-19
- Publication Date
- 2026-03-04
Smart Images

Figure 2026035348000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] While modern generative AI technology and IT tools have provided many benefits, the elderly are currently unable to effectively utilize these tools. In particular, the user interface (UI) is often complex, leaving the elderly confused about how to operate these tools. This problem prevents the elderly from fully enjoying these benefits, creating a digital divide. Therefore, the present invention aims to solve this problem by helping the elderly operate these devices easily. [Means for solving the problem]
[0005] The present invention provides a system that uses AR glasses designed for the elderly to perform image recognition of the user interface of a device operated by the elderly using a camera. The camera means built into the AR glasses captures an image of the user interface and sends the image to a cloud server. The cloud server analyzes the image using a generative artificial intelligence model (generative AI) and generates operation procedures based on the analysis results. The generated operation procedures are provided to the user as audio guidance and AR displays. In addition, the system accepts audio feedback from the user and provides additional operation guidance based on the feedback, thereby achieving interactive support. As a result, even if an elderly person is unsure how to operate a device, they can proceed with the operation while receiving appropriate guidance.
[0006] "Elderly users" refers to users of an age group who require special assistance due to interface complexity and unfamiliarity with operation.
[0007] "AR glasses" refers to a glasses-type device that users can wear to display augmented reality information overlaid on real-world information.
[0008] "Camera means" refers to a device built into the AR glasses for capturing an image of the user interface of a device that is within the user's field of view.
[0009] "Communication means" refers to a device or software that has the function of transmitting images captured by the camera means to a cloud server.
[0010] "Cloud server" refers to a server that provides remote data storage and processing and is accessible via the Internet.
[0011] "Analysis means" refers to a device or software that analyzes image data sent to the cloud server using a generative AI model.
[0012] "Generative artificial intelligence model (generative AI)" refers to a machine learning model that generates new information and guidance based on provided data.
[0013] "Guide means" refers to a device or software that provides the user with the operating procedure generated by the analysis means as audio guidance and AR display.
[0014] "Interactive means" refers to a device or software that accepts a user's voice feedback and generates and provides additional operating guidance based on that feedback.
[0015] "Speech recognition means" means technology or software for converting a user's speech into text data. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] The present invention relates to a support system for elderly people to utilize generative AI and IT tools. This system uses an AR glasses camera to image-recognize the user interface of the device operated by the user, and analyzes the image using a generative AI model to provide operational support using audio guidance and AR displays. The following describes in detail the embodiments of the present invention.
[0038] System Configuration
[0039] The system mainly includes the following components:
[0040] AR glasses: Worn by elderly users, these glasses have built-in cameras, microphones, and speakers.
[0041] Camera means: Built into the AR glasses, it captures the user interface of the device within the user's field of view.
[0042] Communication method: Captured image data is sent to a cloud server.
[0043] Cloud server: Receives image data and has a generative AI model that performs analysis.
[0044] Analysis method: Analyze the user interface in the image using a generative AI model.
[0045] Guidance means: Operating procedures are generated based on the analysis results and presented to the user through audio guidance and AR displays.
[0046] Interaction methods: Receives user voice feedback and generates additional operation instructions based on the feedback.
[0047] Speech recognition means: converts the user's voice feedback into text data.
[0048] Program processing flow
[0049] 1. The user puts on the AR glasses
[0050] The user puts on the AR glasses and turns on the device, which activates the camera, microphone, and other components.
[0051] 2. User Operation
[0052] The user places the device (e.g., smartphone, tablet, or household appliance) they want to operate within their field of view, and the camera captures the user interface of this device in real time.
[0053] 3. Data Transmission
[0054] The terminal transmits the captured image data to a cloud server using a communication means.
[0055] 4. Cloud server analysis
[0056] The server analyzes the received image data using an analysis means and generates device operation instructions using a generative AI model, such as "Tap the Settings app."
[0057] 5. Provision of guides
[0058] The server converts the generated operating procedures into audio files and AR display data and sends them to the terminal.
[0059] Based on this, the device will use voice guidance and AR displays to guide the user on how to operate the device.
[0060] 6. User Feedback
[0061] If the user is confused about an operation, the device will provide voice feedback through the microphone, for example, asking, "What should I do next?"
[0062] 7. Voice Recognition and Further Guidance
[0063] The terminal converts the voice feedback into text using a voice recognition means and transmits it to a cloud server.
[0064] The server analyzes this feedback, generates additional operating instructions, and again uses the guide means to provide operating assistance to the user.
[0065] Specific examples
[0066] Example 1: Smartphone Wi-Fi settings
[0067] The user puts on the AR glasses and projects the smartphone's home screen with the camera. The device sends this image to a cloud server, which then uses generative AI to analyze the image and generate the instruction, "Tap the Settings app." The AR glasses then present this instruction to the user through audio guidance and AR displays. If the user is unsure of the next step, they can ask, "What should I do next?" and the server will continue to generate and provide additional instructions.
[0068] Example 2: Configuring Home Appliances
[0069] The user puts on the AR glasses and places the control panel of a household appliance in their field of view. The device sends an image of this panel to the server, which analyzes it. For example, an instruction such as "Press the power button" is generated and presented to the user through the guidance means. If the user is still unsure about how to operate the device, assistance continues through similar voice feedback and analysis.
[0070] In accordance with the above aspects, the present invention realizes a system that enables elderly people to easily understand and use the operation of complex digital devices.
[0071] The processing flow will be explained below.
[0072] Step 1:
[0073] The user puts on the AR glasses and turns on the device, which activates the camera and microphone and puts them into standby mode.
[0074] Step 2:
[0075] The user brings the device to be operated (e.g., a smartphone or a household appliance) into view, and the camera means captures the user interface of the device in real time.
[0076] Step 3:
[0077] The device compresses and encodes the captured images and sends them to a cloud server with low latency.
[0078] Step 4:
[0079] The server decodes the image data sent and inputs it into a generative artificial intelligence model (generative AI) to begin analysis.
[0080] Step 5:
[0081] A generative AI built into the server analyzes UI elements in the image and identifies the device's current state (e.g., "home screen" or "settings screen").
[0082] Step 6:
[0083] Based on the analysis results, the server generates specific instructions in text format for the next steps, such as "Tap the settings icon."
[0084] Step 7:
[0085] The operating procedures generated by the server are converted into audio files and AR display data and sent to the terminal.
[0086] Step 8:
[0087] Based on the operation guide received by the device from the server, the device presents the user with operation procedures using voice guidance and AR display. For example, the voice guidance may say "Tap the Settings app" and the AR display may show the location of the Settings app.
[0088] Step 9:
[0089] The user operates the device according to the operating instructions. If the user is unsure of an operation, they can ask questions or provide feedback by speaking into the microphone. For example, "What should I do next?"
[0090] Step 10:
[0091] The device captures the user's voice and converts it into text data using voice recognition technology.
[0092] Step 11:
[0093] The terminal transmits the converted text data to the server.
[0094] Step 12:
[0095] The server then re-feeds the user's questions and feedback into the generative AI model to generate an appropriate response, such as "Select your Wi-Fi settings."
[0096] Step 13:
[0097] The server converts the response into an audio file or AR display data and sends it to the device.
[0098] Step 14:
[0099] The device will again provide the response to the user through audio guidance and AR display.
[0100] Step 15:
[0101] The user resumes operation according to further instructions from the server, repeating steps 9 through 14 as necessary.
[0102] Example 1
[0103] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0104] Currently, the difficulty that older people experience when operating complex electronic devices and software is a major problem. Understanding new operation methods and interfaces is a challenge, especially when introducing new technologies and tools, creating technical barriers. To address this issue, assistance systems that can be operated intuitively by older people are needed, and the importance of real-time operation guidance and feedback systems is particularly increasing.
[0105] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0106] In this invention, the server includes a camera, a communication device, and an analysis device that uses a generative AI model. This enables elderly users to operate electronic devices quickly and accurately through real-time, intuitive voice guidance and AR displays. Specifically, the server captures images of the user's operating environment with a camera, sends the images to the server, and analyzes them with a generative AI model to provide appropriate operation guidance. Furthermore, the server converts the user's voice feedback into text for further operational assistance, making it easier to operate complex digital devices.
[0107] "Elderly users" refers to older individuals who are the primary users of the system and who have difficulty using electronic devices and software, especially those that are complex to operate.
[0108] "AR glasses" are headset-type devices that users can wear and use, equipped with augmented reality (AR) technology to display information as an overlay on their field of vision.
[0109] "Photographing means" refers to the camera built into the AR glasses, a device that has the function of capturing an image of the user interface of an electronic device within the user's field of view.
[0110] "Communication means" refers to data transmission devices such as Wi-Fi modules and 4G / 5G communication modules for transmitting captured image data to a remote server.
[0111] A "generative artificial intelligence model" is an AI model used for image recognition and natural language processing, and refers to a program that generates specific operating instructions and guides based on deep learning technology, for example.
[0112] "Analysis means" refers to software and hardware for analyzing the transmitted image data and recognizing the state of the electronic device, the position of the user interface, and the like.
[0113] "Operation guidance means" refers to a mechanism for providing users with operation procedures generated based on the analysis results through audio guidance and AR displays.
[0114] "Interactive means" refers to a device or program that has the function of receiving voice feedback from a user and generating and providing additional operation guidance based on that feedback.
[0115] "Speech recognition means" means the technology or software used to convert a user's voice feedback into text data.
[0116] "Presentation means" refers to a screen display or audio output device that visualizes and presents to the user the operation instructions generated based on the analyzed data.
[0117] This invention relates to a support system that enables elderly people to easily operate complex electronic devices. This system includes a photographing means, a communication means, an analysis means using a generative artificial intelligence model, an operation guidance means, a dialogue means, a voice recognition means, and a presentation means. The following describes in detail an embodiment of the invention.
[0118] Hardware and Software Configuration
[0119] Hardware used
[0120] 1. AR Glasses:
[0121] Built-in camera: A camera for capturing the user interface of an electronic device in the user's field of view.
[0122] Built-in microphone: A microphone for collecting user voice feedback.
[0123] Built-in speaker: A speaker for playing audio guides.
[0124] Communication module: A module that supports Wi-Fi and 4G / 5G communications.
[0125] 2. Cloud Server:
[0126] high performance computer
[0127] Generative artificial intelligence models (e.g., GPT-4 (registered trademark))
[0128] Image Analysis Software
[0129] Software used
[0130] 1. Real-time image capture software: Software that captures images from the AR glasses camera and sends them to a cloud server.
[0131] 2. Speech recognition software: Software for converting user voice feedback into text data.
[0132] 3. Generative AI model: Software that analyzes the transmitted image data and generates appropriate operating instructions.
[0133] 4. Speech synthesis software: Software for converting the generated operating instructions into audio guidance.
[0134] 5. AR display software: Software for overlaying operation guides onto the user's field of view.
[0135] System Operation
[0136] Example 1: Smartphone Wi-Fi settings
[0137] Consider a scenario where a user wears AR glasses and operates the home screen of a smartphone. When the user looks at the smartphone, the camera in the AR glasses captures the home screen, and the image data is sent to a cloud server. The cloud server analyzes this image using a generative AI model and generates an operation instruction such as "Tap the Settings app." This instruction is converted into an audio file and AR display data, which are then sent to the device. The user follows these instructions to perform the operation.
[0138] As a concrete example, the following prompt sentences could be input to a generative AI model:
[0139] "I want to set up Wi-Fi on my smartphone. Can you please tell me the detailed steps to tap the Settings app?"
[0140] Example 2: Home electronics settings
[0141] Consider a scenario where a user wears AR glasses and operates a home electronic device. For example, let's say they want to operate an air conditioner remote control. When the user picks up the remote control, the camera in the AR glasses captures the control panel, and the image is sent to a cloud server. After analysis by the generative AI model, an instruction such as "Press the power button" is generated. This instruction is provided to the user through audio guidance and AR displays.
[0142] As a concrete example, the following prompt sentences could be input to a generative AI model:
[0143] "I want to set up a home appliance. First, please tell me the steps to press the power button."
[0144] As described above, this system helps seniors easily understand and operate complex electronic devices. By providing real-time operation guidance and additional guidance based on user feedback, it is possible to reduce technological barriers and promote the use of digital devices.
[0145] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0146] Step 1:
[0147] The user puts on the AR glasses and turns on the device.
[0148] Specific operation: The user presses and holds the power button on the AR glasses to start the system, which automatically turns on the camera, microphone, speaker, etc.
[0149] Input: User action (pressing the power button)
[0150] Output: Notification of AR glasses startup status and readiness
[0151] Step 2:
[0152] Keep the device the user wants to control within sight.
[0153] Specific Action: A user brings an electronic device, such as a smartphone or consumer electronic device, into their field of view. The camera in the AR glasses captures this user interface.
[0154] Input: Placing the device in view
[0155] Output: Captured image data of the user interface
[0156] Step 3:
[0157] The device transmits the captured image data to a cloud server.
[0158] Specific operation: The device's communication means sends the captured image data to the cloud server via Wi-Fi or 4G / 5G. The user is notified of the progress during the transmission.
[0159] Input: Captured image data
[0160] Output: Notification of completion of data transmission to the cloud server
[0161] Step 4:
[0162] The server analyzes the received image data.
[0163] Specific operation: The server begins analyzing the received image data using analytical means. A generative AI model (e.g., GPT-4) is used for the analysis to recognize the device state and interface.
[0164] Input: Image data sent
[0165] Output: Parsed data (current state of the device and possible actions)
[0166] Step 5:
[0167] The server generates operation instructions using the generative AI model.
[0168] Specific actions: The generative AI model generates appropriate instructions based on the analysis results. For example, it creates specific instructions such as "Tap the Settings app."
[0169] Input: Parsed data
[0170] Output: Generated operating instructions
[0171] Step 6:
[0172] The operating procedures generated by the server are converted into audio files and AR display data.
[0173] Specific operation: The server converts the generated operation instructions into a voice file and AR display data using voice synthesis software and AR display software.
[0174] Input: Generated operation instructions
[0175] Output: Audio file and AR display data
[0176] Step 7:
[0177] The server sends the generated audio files and AR display data to the terminal.
[0178] Specific operation: The server sends the audio file and AR display data to the device using a communication method. The user is notified of the transmission status.
[0179] Input: Audio files and AR display data
[0180] Output: Notification of completion of transmission to the terminal
[0181] Step 8:
[0182] The device will guide the user on how to operate the device through voice guidance and AR displays.
[0183] Specific operation: The device plays the received audio file through the speaker and displays an overlay of operation instructions on the AR glasses display, such as "Tap the Settings app."
[0184] Input: Audio files and AR display data
[0185] Output: User guide
[0186] Step 9:
[0187] If the user has difficulty operating the device, voice feedback is provided.
[0188] Specific Action: If a user is unsure about an action, they can speak into the microphone and ask a question like, "What should I do next?" The microphone will capture their voice.
[0189] Input: User's voice feedback
[0190] Output: Captured audio data
[0191] Step 10:
[0192] The terminal converts the voice feedback into text using a voice recognition means and transmits it to a cloud server.
[0193] Specific operation: The device converts the voice data into text data using voice recognition software and sends the text data to the cloud server.
[0194] Input: Captured audio data
[0195] Output: Text data and notification of completion of transmission to the cloud server
[0196] Step 11:
[0197] The server analyzes the text data and generates additional operating instructions.
[0198] Specific operation: The server analyzes the received text data and generates additional operation instructions based on the user's question, such as "Next, select Wi-Fi."
[0199] Input: Text data
[0200] Output: Additional operating instructions
[0201] Step 12:
[0202] The server transmits the generated additional operation instructions to the terminal, and the terminal provides the user with additional guidance.
[0203] Specific operation: The server converts additional operation instructions into an audio file and AR display data and sends them to the terminal, which then provides guidance to the user based on this.
[0204] Input: Additional operating instructions
[0205] Output: Operation guide with additional audio files and AR display data
[0206] The above is the flow of processing in the program for this system and the specific operations of each.
[0207] (Application example 1)
[0208] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0209] When elderly people make electronic payments using smartphones and other digital devices, the operation is complicated and difficult for them to use. In addition, there is a high possibility that elderly people will make operational errors or face security risks, so an efficient method to support them is needed.
[0210] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0211] In this invention, the server includes an analysis means, an operation guidance means, and a dialogue means, which allows elderly users to be guided through the transaction procedure through the AR glasses when making electronic payments using their smartphones, and provides specific operation instructions according to the progress of the transaction.
[0212] "Elderly users" refer to users who are older and unfamiliar with operating digital devices.
[0213] "Wearable AR glasses" are glasses-type devices that users wear on their heads to display augmented reality.
[0214] A "user interface" refers to the part of a device that a user directly operates or inputs, such as a screen or operation panel.
[0215] "Imaging means" refers to the function of capturing images using a camera or sensor.
[0216] "Data communication means" refers to the function of transmitting captured data to other devices or servers.
[0217] A "cloud computing server" refers to a remote server used to store and analyze data over the Internet.
[0218] A "generative artificial intelligence model" is an artificial intelligence that analyzes input and data from users and generates appropriate responses and instructions.
[0219] "Analysis means" refers to a function that analyzes the captured image data and extracts information about the user's operations.
[0220] "Operation guide means" refers to a function that presents specific operation methods to the user based on the analysis results.
[0221] "Augmented reality display" refers to a technology that adds virtual information to the real field of view and displays it.
[0222] "Interaction" refers to the ability to interact with the user and provide additional guidance based on their feedback.
[0223] "Electronic payment support means" refers to a function that supports operations during electronic payment and guides transaction procedures.
[0224] "Transaction guide means" refers to a function that provides specific operating instructions, such as scanning a QR code (registered trademark) or tapping a payment button, depending on the progress of the transaction.
[0225] The system for realizing this invention is designed to help elderly users smoothly make electronic payments. The system is realized by the following components and operations.
[0226] System Configuration
[0227] The system mainly includes the following components:
[0228] 1. AR glasses: A device worn by elderly users that has a built-in camera, microphone, and speaker.
[0229] 2. Imaging means: The camera built into the AR glasses captures the device's user interface in real time as it comes into the user's field of view.
[0230] 3. Data communication means: Transmit the captured image data to a cloud computing server.
[0231] 4. Cloud computing server: Analyzes the received image data and generates operation guides using a generative AI model.
[0232] 5. Analysis method: Analyze the user interface in the image on the cloud server and identify the next operation to be performed.
[0233] 6. Operation guide means: Based on the analysis results, operation procedures are generated as audio guides and AR displays and provided to the user.
[0234] 7. Interaction: Receives voice feedback from the user and provides additional guidance based on that.
[0235] 8. Speech recognition means: converts the user's voice feedback into text data.
[0236] 9. Electronic payment support measures: Provide specific operating procedures for electronic payments.
[0237] 10. Transaction Guidance: Guidance on transaction operations such as scanning a QR code or tapping a payment button.
[0238] How it works
[0239] 1. Function of the imaging means: When a user wears the AR glasses and tries to operate a smartphone, the camera captures the smartphone screen and obtains image data in real time.
[0240] 2. Function of data communication means: The acquired image data is transmitted to the cloud computing server via data communication means (e.g., Wi-Fi or mobile network).
[0241] 3. Function of the analysis tool: A generative AI model on a cloud server analyzes image data, for example, identifying QR codes or specific buttons.
[0242] 4. Operation guidance function: Based on the analysis results, the cloud server generates specific operation instructions, such as "Scan the QR code" or "Tap the payment button."
[0243] 5. Audio guide and AR display: The generated operating procedure is provided to the user as an audio guide and an augmented reality display. The next operation the user should perform is presented visually and audibly.
[0244] 6. Interaction function: If the user is unsure of what to do, they can provide voice feedback (e.g., "Which button should I press next?"). This feedback is sent to the server via the microphone.
[0245] 7. Voice recognition and additional guidance: Voice feedback is converted into text data by a voice recognition tool and analyzed again by the cloud server. Additional specific operating instructions are then generated and provided again as voice guidance or AR displays.
[0246] Specific examples
[0247] Example 1: Payment processing for an electronic payment app
[0248] The user puts on the AR glasses and opens the electronic payment app. The smartphone screen is captured and analyzed by the cloud server. The server generates an instruction to "scan the QR code" and displays it as an audio guide and AR display through the AR glasses. When the user follows the instructions, the next instruction is displayed: "tap the payment button."
[0249] Prompt: A user is looking at an electronic payment app. Guide them through the next steps in the payment process.
[0250] Instructions for the generative AI model: "Scan the QR code" or "Tap the payment button"
[0251] Example 2: Feedback and additional guidance
[0252] The user is confused and asks, "Which button should I press next?" This voice input is converted into text data and sent to a cloud server. The server analyzes the feedback and generates additional instructions as the next step: "Tap the back button and reselect your payment method."
[0253] This invention enables elderly people to easily understand how to operate complex digital devices and use them safely.
[0254] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0255] Step 1:
[0256] A user puts on the AR glasses and attempts to operate a smartphone. The user launches an electronic payment app, and the camera in the AR glasses captures this screen in real time.
[0257] Input: The smartphone screen operated by the user
[0258] Output: Real-time captured image data
[0259] Specific operation: The camera in the AR glasses captures an image of the user interface and acquires it as image data.
[0260] Step 2:
[0261] The terminal transmits the acquired image data to a cloud computing server via a communication means, with the data transfer occurring via an internet connection.
[0262] Input: Image data captured in real time
[0263] Output: Image data sent to a cloud computing server
[0264] Specific operation: The communication module in the AR glasses transmits image data using Wi-Fi or a mobile network.
[0265] Step 3:
[0266] The server analyzes the received image data using an analytical method and uses a generative AI model to determine the next step to take, such as whether the user should scan a QR code or tap a payment button.
[0267] Input: Image data sent to a cloud computing server
[0268] Output: Specific operation procedures as analysis results
[0269] Specific operation: The generative artificial intelligence model analyzes the image data and generates appropriate operating instructions based on the prompt text.
[0270] Step 4:
[0271] The server generates the operating procedures as audio guidance and AR display and sends them to the terminal.
[0272] Input: Specific operation procedures as analysis results
[0273] Output: Audio guide and AR display data
[0274] Specific operation: The operation guide generation module in the server converts the instructions into audio files and AR display data.
[0275] Step 5:
[0276] The device uses the received voice guidance and AR display to guide the user on how to operate the device.
[0277] Input: Audio guide and AR display data
[0278] Output: Audio guidance and AR display to the user
[0279] What it does: Uses the speakers and display devices of the AR glasses to provide audio and visual guidance.
[0280] Step 6:
[0281] If the user is confused about an operation, they can receive voice feedback, such as asking, "Which button should I press next?" This voice feedback is collected through the microphone.
[0282] Input: User's voice feedback
[0283] Output: Collected audio data
[0284] Specific operation: The microphone inside the AR glasses captures sound and saves it as audio data.
[0285] Step 7:
[0286] The terminal converts the voice feedback into text using a voice recognition means and transmits it to a cloud server.
[0287] Input: Collected audio data
[0288] Output: Feedback data converted to text
[0289] Specific operation: Uses voice recognition technology to convert voice data into text data.
[0290] Step 8:
[0291] The server analyzes the text feedback and generates additional operating instructions, which are then provided to the user again as audio guidance and AR displays.
[0292] Input: Feedback data converted to text
[0293] Output: Audio guide and AR display data as additional operating instructions
[0294] Specific behavior: The generative AI model analyzes the feedback, generates new operating instructions, and presents them to the user again.
[0295] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0296] The present invention combines an emotion engine with a support system for elderly users to utilize generative AI. This system uses an AR glasses camera to image-recognize the user interface of the device operated by the user, and analyzes the image using a generative AI model to provide operational support using audio guidance and AR displays, as well as flexible guidance by recognizing the user's emotions. The following describes in detail the embodiments of the present invention.
[0297] System Configuration
[0298] The system mainly includes the following components:
[0299] AR glasses: Worn by elderly users, these glasses have built-in cameras, microphones, and speakers.
[0300] Camera means: Built into the AR glasses, it captures the user interface of the device within the user's field of view.
[0301] Communication means: Captured image data and audio data are sent to a cloud server.
[0302] Cloud server: Receives image data and audio data and has a generative AI model that analyzes them.
[0303] Analysis method: A generative AI model is used to analyze the user interface and audio data in the image.
[0304] Guidance means: Based on the analysis results, operating procedures are generated and presented to the user through audio guidance and AR displays.
[0305] Interaction methods: Accepts user voice feedback and provides additional guidance based on the feedback.
[0306] Speech recognition means: converts the user's voice feedback into text data.
[0307] Emotion Engine: Evaluates emotions by analyzing the user's tone of voice, facial expressions, and posture.
[0308] Program processing flow
[0309] 1. The user puts on the AR glasses
[0310] The user puts on the AR glasses and turns them on, which activates the camera and microphone and puts them into standby mode.
[0311] 2. User Actions
[0312] The user brings the device to be operated (e.g., a smartphone or a household appliance) into view, and the camera means captures the user interface of this device in real time.
[0313] 3. Data Transmission
[0314] The image and audio data captured by the device is compressed and encoded and sent to a cloud server with low latency.
[0315] 4. Cloud server analysis
[0316] The image and audio data received by the server is analyzed by an analysis means, and the user interface and audio are analyzed using a generative artificial intelligence model (generative AI).
[0317] The emotion engine analyzes voice tone, facial expressions, and posture to assess the user's emotions.
[0318] 5. Providing operation guides
[0319] Based on the analysis results and emotion evaluation, the server converts the specific next steps into text format, converts it into an audio file or AR display data, and sends it to the device.
[0320] Based on this, the device will present the user with operation instructions using voice guidance and AR display. For example, the voice guidance will say "Tap the Settings app" and the AR display will show the location of the Settings app.
[0321] 6. User Feedback
[0322] The user operates the device by following the instructions. If they are unsure of how to operate the device, they can ask questions or provide feedback by speaking into the microphone. For example, "What should I do next?"
[0323] 7. Speech Recognition and Emotion Assessment
[0324] The device captures the user's voice feedback and converts it into text data using a speech recognition tool, while the emotion engine evaluates the user's emotions.
[0325] 8. Submitting Feedback
[0326] The device sends the converted text data and emotional information to a cloud server.
[0327] 9. Response generation by cloud server
[0328] The server then re-feeds the user's questions and feedback into the generative AI model to generate an appropriate response, such as "Select your Wi-Fi settings."
[0329] Based on information from the emotion engine, the content and presentation of the response can be flexibly adjusted, for example, if the user is annoyed, a calmer tone of response can be generated.
[0330] 10. Providing a Response
[0331] The server converts the response into an audio file or AR display data and sends it to the device.
[0332] The device will again provide the response to the user through audio guidance and AR display.
[0333] Specific examples
[0334] Example 1: Smartphone Wi-Fi settings
[0335] The user puts on the AR glasses and projects the smartphone's home screen with the camera. The device sends this image to a cloud server, which then uses generative AI to analyze the image and generate the instruction, "Tap the Settings app." The AR glasses then present this instruction to the user through voice guidance and AR display. If the user is unsure of the next step, they can ask aloud, "What should I do next?" If the emotion engine detects the user's frustration along with this question, the server generates and provides more detailed instructions in a gentler tone.
[0336] Example 2: Configuring Home Appliances
[0337] The user puts on the AR glasses and brings the control panel of a household appliance into view. The device sends an image of this panel to the server, which analyzes it. For example, an instruction such as "Press the power button" is generated and presented to the user through a guide means. If the user is unsure and gives feedback such as "This doesn't work," the emotion engine will detect anxiety and the server will generate an instruction in an encouraging tone such as "Don't worry. Next, try..."
[0338] In accordance with the above aspects, the present invention realizes a system that enables elderly people to easily understand and use the operation of complex digital devices, and provides flexible support that takes emotions into consideration.
[0339] The processing flow will be explained below.
[0340] Step 1:
[0341] The user puts on the AR glasses, which activates the camera and microphone of the AR glasses and puts them into standby mode.
[0342] Step 2:
[0343] The user brings the device to be operated (e.g., a smartphone or a home appliance) into view. The device uses a camera to capture the device's user interface in real time.
[0344] Step 3:
[0345] The device compresses and encodes the captured image data and user voice data and sends them to the cloud server.
[0346] Step 4:
[0347] The server decodes the received image data, and the analysis means analyzes the user interface elements in the image using a generative artificial intelligence model (generative AI).
[0348] Step 5:
[0349] The server analyzes the received voice data and converts it into text using a voice recognition means.
[0350] Step 6:
[0351] The server uses the generation AI to generate the next steps based on the analysis results, creating specific instructions such as "tap the Settings app."
[0352] Step 7:
[0353] The server converts the generated operating procedures into audio files and AR display data and sends them to the terminal.
[0354] Step 8:
[0355] Based on the operation guide received by the device from the server, the device presents the user with operation procedures using voice guidance and AR display. For example, the voice guidance may say "Tap the Settings app" and the AR display may show the location of the Settings app.
[0356] Step 9:
[0357] The user operates the device according to the operating instructions. If the user is unsure of how to operate the device, they can ask questions into the microphone, such as "What should I do next?"
[0358] Step 10:
[0359] The device captures the user's voice feedback and converts it into text data using voice recognition technology, while the emotion engine analyzes the user's voice tone, facial expressions, and posture to make an emotional assessment.
[0360] Step 11:
[0361] The device sends the converted text data and emotional information to a cloud server.
[0362] Step 12:
[0363] The server re-inputs the user's questions and feedback into the AI model to generate an appropriate response. Based on the evaluation of the emotion engine, the model generates instructions with a tone and content that corresponds to the user's emotions. For example, if the user is showing signs of irritation, the model generates instructions in a calm tone such as "Don't worry. Next, please select your Wi-Fi settings."
[0364] Step 13:
[0365] The server converts the response into an audio file or AR display data and sends it to the device.
[0366] Step 14:
[0367] The device will again provide the response to the user through audio guidance and AR display.
[0368] Step 15:
[0369] The user resumes operation following any further instructions from the server, repeating steps 9 through 14 as necessary.
[0370] Example 2
[0371] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0372] When elderly people operate complex digital devices, they often find it difficult to understand the operating procedures, and emotions such as anxiety and frustration during operation can hinder their operation. To solve this problem, it is necessary to provide appropriate operating instructions that take into account the user's emotions.
[0373] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0374] In this invention, the server includes an interaction means for receiving voice feedback from the user and providing additional operation guidance based on the feedback, a voice recognition means for converting the user's voice feedback into text using voice recognition technology, and an emotion evaluation means for analyzing the voice tone, facial expression, and posture to evaluate the user's emotions, thereby making it possible to provide flexible and appropriate operation guidance that takes the user's emotions into consideration.
[0375] "Elderly users" refer to users who are likely to find it difficult to operate digital devices due to their age.
[0376] "AR glasses" refers to a glasses-type device that uses augmented reality technology to overlay information on the real world.
[0377] "Capture means" refers to a camera or other capture device used to capture images of objects or interfaces within the user's field of view.
[0378] "Communication means" refers to an internet connection or wireless communication means for sending and receiving image and audio data between the terminal and the cloud server.
[0379] A "generative artificial intelligence model" refers to a machine learning model designed to perform a specific task (e.g., image recognition or speech analysis).
[0380] The "analysis means" refers to a digital processor that analyzes the acquired image and audio data, recognizes the user interface, and generates operation procedures.
[0381] "Guidance means" refers to an audio output device or an augmented reality display device that notifies the user of the operating procedure generated based on the analysis results.
[0382] "Interaction means" refers to the part of the system that accepts voice or other operational input from the user and generates additional operational instructions based on that input.
[0383] "Speech recognition means" refers to technology or equipment for converting voice data obtained from a user into text data.
[0384] "Emotion assessment means" refers to technology or devices for analyzing a user's tone of voice, facial expression, posture, etc. to assess the user's emotional state.
[0385] This system aims to assist elderly users in operating complex digital devices, and uses AR glasses to capture the user interface of the target device within the user's field of view. Specifically, a camera means built into the AR glasses captures images of the user device in real time and transmits them to a cloud server via a communication means.
[0386] Hardware and software used
[0387] AR glasses: Devices that display information using augmented reality technology and have built-in cameras, microphones, and speakers.
[0388] Camera means: Capable of capturing the user interface of a device within the user's field of view.
[0389] Communication means: Includes internet connection and wireless communication means for transmitting captured image and audio data to a cloud server.
[0390] Cloud server: Receives image and audio data and analyzes them using generative artificial intelligence models.
[0391] Analysis means: A generative artificial intelligence model in the cloud server is used to analyze the user interface and audio data in the image.
[0392] Guidance means: Based on the analysis results, operating procedures are generated and presented to the user through audio guidance and augmented reality displays.
[0393] Interaction methods: Accepts user voice feedback and provides additional guidance.
[0394] Speech recognition means: Uses technology to convert the user's voice feedback into text data.
[0395] Emotion assessment means: It has technology to assess emotions by analyzing the user's tone of voice, facial expressions, and posture.
[0396] System operation explanation
[0397] When a user puts on the AR glasses and turns them on, the camera and microphone are activated and in standby mode. When the user brings the device to be operated (e.g., a smartphone, tablet, or household appliance) into view, the camera captures the device's user interface. The captured image and audio data are sent to a cloud server via a communication means.
[0398] The cloud server analyzes the received image and audio data using an analysis means and uses a generative artificial intelligence model to analyze the user interface and audio. The emotion evaluation means also analyzes the voice tone, facial expression, and posture to evaluate the user's emotions. Based on the analysis results and the emotion evaluation, specific next steps are generated, converted into audio files and augmented reality display data, and then sent to the device.
[0399] The device provides the user with operational instructions via voice guidance and augmented reality display. For example, the voice guidance might say, "Tap the Settings app," and the location of the Settings app is highlighted on the AR glasses' display. If the user is unsure of the next step, they can ask into the microphone, "What should I do next?". The emotion assessment means analyzes the user's emotions along with this question, and the cloud server generates and provides detailed, gentle-toned instructions.
[0400] Specific examples
[0401] Example 1: Smartphone Wi-Fi settings
[0402] The user puts on the AR glasses and projects the smartphone's home screen with the camera. The device sends this image to a cloud server, which then uses generative AI to analyze the image and generate the instruction, "Tap the Settings app." The AR glasses then present this instruction to the user through voice guidance and an augmented reality display. If the user is unsure of the next step, they can ask aloud, "What should I do next?" If the emotion engine detects the user's frustration along with this question, the server generates and provides more detailed instructions in a gentler tone.
[0403] Example 2: Configuring Home Appliances
[0404] The user puts on the AR glasses and brings the control panel of a household appliance into view. The device sends an image of this panel to the server, which analyzes it. For example, an instruction such as "Press the power button" is generated and presented to the user through a guide means. If the user is unsure and gives feedback such as "This doesn't work," the emotion engine will detect anxiety and the server will generate an instruction in an encouraging tone such as "Don't worry. Next, try..."
[0405] Prompt Sentence Examples
[0406] "Please provide audio guidance and AR displays to guide users through the steps they need to set up Wi-Fi on their smartphones."
[0407] The above system aims to use a generative AI model and an emotion engine to enable elderly users to easily operate complex digital devices, and to provide flexible assistance that takes into account emotions during operation.
[0408] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0409] Step 1:
[0410] The user puts on the AR glasses.
[0411] Input: The user puts on the AR glasses and turns them on.
[0412] Operation: The camera and microphone of the AR glasses are activated and enter standby mode.
[0413] Output: The camera and microphone are ready to use.
[0414] Step 2:
[0415] Keep the device the user is controlling within sight.
[0416] Input: The user brings the device they want to control (e.g., a smartphone or household appliance) into view.
[0417] Operation: A camera means captures the device's user interface in real time.
[0418] Output: Captured image and audio data.
[0419] Step 3:
[0420] Send the data to a cloud server.
[0421] Input: Captured image and audio data.
[0422] Operation: The device compresses and encodes image and audio data and sends them to the cloud server using a communication method.
[0423] Output: The cloud server receives the captured image and audio data.
[0424] Step 4:
[0425] A cloud server analyzes the image and audio data.
[0426] Input: Image data and audio data received by the cloud server.
[0427] Operation: The analysis means (generative AI model) analyzes image data and recognizes the user interface. The emotion evaluation means analyzes voice tone, facial expressions, and posture to evaluate the user's emotions.
[0428] Output: User interface analysis results and user emotion evaluation results.
[0429] Step 5:
[0430] Generate operation guide.
[0431] Input: User interface analysis results and user emotion evaluation results.
[0432] How it works: The cloud server uses the generative AI model to create specific operating procedures in text format based on the analysis results and emotional evaluation.
[0433] Output: Generated operating instructions text.
[0434] Step 6:
[0435] The device provides operation guidance.
[0436] Input: The generated operating instructions text.
[0437] How it works: The device converts the instruction text into an audio file and augmented reality display data, and presents it to the user. Specifically, the audio guide says "Tap the Settings app," and the AR display highlights the location of the Settings app.
[0438] Output: Audio guide and augmented reality display presented to the user.
[0439] Step 7:
[0440] Process user feedback.
[0441] Input: The user interacts with the device through instructions and uses voice to ask questions or provide feedback, for example, "What should I do next?"
[0442] What it does: The device captures the user's voice feedback.
[0443] Output: Captured audio feedback data.
[0444] Step 8:
[0445] Performs voice recognition and emotion assessment.
[0446] Input: Captured audio feedback data.
[0447] Operation: The device uses a voice recognition means to convert voice data into text, and an emotion assessment means analyzes the user's voice tone and facial expressions to assess their emotions.
[0448] Output: Feedback and sentiment evaluation results converted into text.
[0449] Step 9:
[0450] Send the feedback to a cloud server.
[0451] Input: Feedback and sentiment rating results converted to text.
[0452] Operation: The device sends text data and emotion information to the cloud server using a communication method.
[0453] Output: The cloud server receives the feedback data and emotion information.
[0454] Step 10:
[0455] The cloud server generates a response.
[0456] Input: Feedback data and emotional information.
[0457] How it works: The server uses a generative AI model to generate appropriate responses based on user feedback, and flexibly adjusts the content and presentation of the response based on emotional assessment, for example, generating a calmer tone of response if the user is annoyed.
[0458] Output: The generated response text.
[0459] Step 11:
[0460] The terminal provides a response.
[0461] Input: The generated response text.
[0462] How it works: The device converts the response text into an audio file and augmented reality display data, and presents them to the user. Specifically, the audio prompt says "Select Wi-Fi settings" and the corresponding part is highlighted in the AR display.
[0463] Output: Audio guide and augmented reality display presented to the user.
[0464] Through the above processing steps, this system can provide appropriate and flexible support for operating digital devices while taking into account the user's emotions.
[0465] (Application example 2)
[0466] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0467] Elderly users face challenges when operating complex digital devices or selecting products in physical stores, such as difficulty obtaining and understanding information about operation and purchasing. Additionally, elderly users often experience anxiety and frustration while operating devices, which can be difficult to deal with appropriately. This can make it difficult for elderly users to smoothly navigate digital devices and physical stores.
[0468] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0469] In this invention, the server includes a camera, a communication unit, an analysis unit using a generative AI model, an emotion evaluation unit, a guide unit, a dialogue unit, and a voice recognition unit. This allows elderly users to easily obtain and understand product information on digital devices and in physical stores using AR glasses, and also allows them to receive flexible operation guidance that takes the user's emotions into consideration.
[0470] "AR glasses" are devices worn by elderly users that can overlay digital information onto their field of vision.
[0471] "Camera means" refers to a device for capturing an image of the user interface of a device within the user's field of view.
[0472] "Communication means" is a function for transmitting captured image data to a cloud server.
[0473] A "generative AI model" is an artificial intelligence model that analyzes received image data and generates appropriate operating procedures.
[0474] "Analysis means" is a function for analyzing transmitted images using a generative AI model.
[0475] The "guide means" is a function that provides users with operating procedures generated based on the analysis results through audio guidance and AR displays.
[0476] The "emotion evaluation means" is a function for analyzing the user's voice and facial expressions and evaluating the user's emotional state.
[0477] "Interactive means" is a function that accepts voice feedback from the user and provides additional operational guidance and emotion-sensitive guidance based on that feedback.
[0478] "Voice recognition means" is a function for converting the user's voice feedback into text data.
[0479] This invention is a system that makes it easier for elderly users to select products using complex digital devices or in physical stores. The system includes AR glasses worn by the elderly user, a camera means, a communication means, an analysis means having a generative AI model, an emotion evaluation means, a guidance means, a dialogue means, and a voice recognition means.
[0480] System Configuration
[0481] 1. AR Glasses:
[0482] This device is worn by the user and has a built-in camera, microphone, and speaker, allowing digital information to be displayed overlaid on the user's field of vision when operating a digital device or selecting products in a physical store.
[0483] 2. Camera means:
[0484] It is built into AR glasses and captures images of the device's user interface and products in physical stores within the user's field of view.
[0485] 3. Means of communication:
[0486] The captured image and audio data are sent to a cloud server using a secure, low-latency protocol.
[0487] 4. Analytical means with generative AI models:
[0488] The image data received by the server is analyzed and appropriate operating procedures and product information are generated. The generative AI model learns from a large data set and provides highly accurate analysis results.
[0489] 5. Emotional assessment measures:
[0490] It analyzes the user's tone of voice, facial expressions, and posture to assess the user's emotional state, allowing it to process user feedback appropriately and provide emotionally sensitive assistance.
[0491] 6. Guide means:
[0492] Based on the analysis results, the system generates specific next steps and product information, and provides these instructions to the user in the form of audio guidance or AR displays.
[0493] 7. Means of interaction:
[0494] If the user is unsure of a procedure or needs additional information, the system accepts voice feedback and provides additional operating instructions or emotionally sensitive guidance based on the feedback.
[0495] 8. Voice Recognition Methods:
[0496] The user's voice feedback is converted into text data, which is further analyzed by generative AI models and sentiment assessment methods.
[0497] Example
[0498] Example 1: Smartphone settings
[0499] The user puts on the AR glasses and projects the smartphone's home screen onto the camera. The cloud server analyzes this image using a generative AI model and provides the instruction, "Tap the Settings app." The AR glasses present this instruction to the user through voice guidance and AR displays. If the user is unsure of the next step, they can ask aloud, "What should I do next?" If the emotion engine detects the user's frustration along with this question, the server generates and provides more detailed instructions in a gentler tone.
[0500] Example 2: Product selection in a store
[0501] A user enters a supermarket and wants to look at a particular shelf. The AR glasses send an image of the shelf to a cloud server. The cloud server analyzes the image using a generative AI model and provides a voice prompt saying, "This product is rich in probiotics." If the user's emotions are unstable, the emotion assessment mechanism detects this and provides instructions in an encouraging tone, such as, "Don't worry, check out this product next."
[0502] Prompt example
[0503] "What should I do next?"
[0504] "Is this the right button to press?"
[0505] As described above, this system allows elderly users to easily operate digital devices and select products in physical stores. Furthermore, the emotional support allows them to use the system with peace of mind.
[0506] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0507] Step 1:
[0508] A user wears AR glasses and places a digital device or product in their field of view. The camera built into the AR glasses captures an image of the user interface or product. This image data becomes the input for the system. The camera means performs real-time image capture.
[0509] Step 2:
[0510] The device sends image data captured by the camera to the cloud server via a communication means. At this time, the image data is compressed, encoded, and sent with low latency. Data transmission is performed safely and quickly using a communication protocol.
[0511] Step 3:
[0512] The server analyzes the received image data using a generative AI model. This analysis method identifies the user interface and product attributes in the image and generates operation procedures and product information as the analysis results. The input is image data, and the output is the analysis results.
[0513] Step 4:
[0514] Based on the analysis results, the server generates the next operation procedure and product information, which the guide means provides to the terminal in the form of audio guide and AR display. The guide means converts the generated text into an audio file and AR display data.
[0515] Step 5:
[0516] The user operates the device and checks the product according to the presented operating instructions and product information. If the user is unsure of the next step, voice feedback is provided through the device's microphone. A prompt such as "What should I do next?" is input.
[0517] Step 6:
[0518] The device captures the user's voice feedback and converts it into text data using a voice recognition means. At the same time, the emotion assessment means analyzes the user's voice tone and facial expressions to evaluate the user's emotional state. The input is voice feedback, and the output is text data and emotion information.
[0519] Step 7:
[0520] The text data and emotional information converted by the speech recognition means are sent to a cloud server. The server then inputs this data back into the generative AI model to generate an appropriate response. The content and presentation of the response are flexibly adjusted based on the analysis results of the emotional evaluation means.
[0521] Step 8:
[0522] The server converts the generated response into an audio file and AR display data and sends them to the terminal. Based on this, the guide means provides the user with a response again in the form of audio guidance and AR display. For example, an instruction such as "Please select Wi-Fi settings" is generated.
[0523] This series of steps allows elderly users to smoothly operate digital devices and select products in physical stores.
[0524] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0525] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0526] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0527] [Second embodiment]
[0528] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0529] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0530] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0531] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0532] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0533] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0534] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0535] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0536] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0537] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0538] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0539] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0540] The present invention relates to a support system for elderly people to utilize generative AI and IT tools. This system uses an AR glasses camera to image-recognize the user interface of the device operated by the user, and analyzes the image using a generative AI model to provide operational support using audio guidance and AR displays. The following describes in detail the embodiments of the present invention.
[0541] System Configuration
[0542] The system mainly includes the following components:
[0543] AR glasses: Worn by elderly users, these glasses have built-in cameras, microphones, and speakers.
[0544] Camera means: Built into the AR glasses, it captures the user interface of the device within the user's field of view.
[0545] Communication method: Captured image data is sent to a cloud server.
[0546] Cloud server: Receives image data and has a generative AI model that performs analysis.
[0547] Analysis method: Analyze the user interface in the image using a generative AI model.
[0548] Guidance means: Operating procedures are generated based on the analysis results and presented to the user through audio guidance and AR displays.
[0549] Interaction methods: Receives user voice feedback and generates additional operation instructions based on the feedback.
[0550] Speech recognition means: converts the user's voice feedback into text data.
[0551] Program processing flow
[0552] 1. The user puts on the AR glasses
[0553] The user puts on the AR glasses and turns on the device, which activates the camera, microphone, and other components.
[0554] 2. User Operation
[0555] The user places the device (e.g., smartphone, tablet, or household appliance) they want to operate within their field of view, and the camera captures the user interface of this device in real time.
[0556] 3. Data Transmission
[0557] The terminal transmits the captured image data to a cloud server using a communication means.
[0558] 4. Cloud server analysis
[0559] The server analyzes the received image data using an analysis means and generates device operation instructions using a generative AI model, such as "Tap the Settings app."
[0560] 5. Provision of guides
[0561] The server converts the generated operating procedures into audio files and AR display data and sends them to the terminal.
[0562] Based on this, the device will use voice guidance and AR displays to guide the user on how to operate the device.
[0563] 6. User Feedback
[0564] If the user is confused about an operation, the device will provide voice feedback through the microphone, for example, asking, "What should I do next?"
[0565] 7. Voice Recognition and Further Guidance
[0566] The terminal converts the voice feedback into text using a voice recognition means and transmits it to a cloud server.
[0567] The server analyzes this feedback, generates additional operating instructions, and again uses the guide means to provide operating assistance to the user.
[0568] Specific examples
[0569] Example 1: Smartphone Wi-Fi settings
[0570] The user puts on the AR glasses and projects the smartphone's home screen with the camera. The device sends this image to a cloud server, which then uses generative AI to analyze the image and generate the instruction, "Tap the Settings app." The AR glasses then present this instruction to the user through audio guidance and AR displays. If the user is unsure of the next step, they can ask, "What should I do next?" and the server will continue to generate and provide additional instructions.
[0571] Example 2: Configuring Home Appliances
[0572] The user puts on the AR glasses and places the control panel of a household appliance in their field of view. The device sends an image of this panel to the server, which analyzes it. For example, an instruction such as "Press the power button" is generated and presented to the user through the guidance means. If the user is still unsure about how to operate the device, assistance continues through similar voice feedback and analysis.
[0573] In accordance with the above aspects, the present invention realizes a system that enables elderly people to easily understand and use the operation of complex digital devices.
[0574] The processing flow will be explained below.
[0575] Step 1:
[0576] The user puts on the AR glasses and turns on the device, which activates the camera and microphone and puts them into standby mode.
[0577] Step 2:
[0578] The user brings the device to be operated (e.g., a smartphone or a household appliance) into view, and the camera means captures the user interface of the device in real time.
[0579] Step 3:
[0580] The device compresses and encodes the captured images and sends them to a cloud server with low latency.
[0581] Step 4:
[0582] The server decodes the image data sent and inputs it into a generative artificial intelligence model (generative AI) to begin analysis.
[0583] Step 5:
[0584] A generative AI built into the server analyzes UI elements in the image and identifies the device's current state (e.g., "home screen" or "settings screen").
[0585] Step 6:
[0586] Based on the analysis results, the server generates specific instructions in text format for the next steps, such as "Tap the settings icon."
[0587] Step 7:
[0588] The operating procedures generated by the server are converted into audio files and AR display data and sent to the terminal.
[0589] Step 8:
[0590] Based on the operation guide received by the device from the server, the device presents the user with operation procedures using voice guidance and AR display. For example, the voice guidance may say "Tap the Settings app" and the AR display may show the location of the Settings app.
[0591] Step 9:
[0592] The user operates the device according to the operating instructions. If the user is unsure of an operation, they can ask questions or provide feedback by speaking into the microphone. For example, "What should I do next?"
[0593] Step 10:
[0594] The device captures the user's voice and converts it into text data using voice recognition technology.
[0595] Step 11:
[0596] The terminal transmits the converted text data to the server.
[0597] Step 12:
[0598] The server then re-feeds the user's questions and feedback into the generative AI model to generate an appropriate response, such as "Select your Wi-Fi settings."
[0599] Step 13:
[0600] The server converts the response into an audio file or AR display data and sends it to the device.
[0601] Step 14:
[0602] The device will again provide the response to the user through audio guidance and AR display.
[0603] Step 15:
[0604] The user resumes operation according to further instructions from the server, repeating steps 9 through 14 as necessary.
[0605] Example 1
[0606] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0607] Currently, the difficulty that older people experience when operating complex electronic devices and software is a major problem. Understanding new operation methods and interfaces is a challenge, especially when introducing new technologies and tools, creating technical barriers. To address this issue, assistance systems that can be operated intuitively by older people are needed, and the importance of real-time operation guidance and feedback systems is particularly increasing.
[0608] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0609] In this invention, the server includes a camera, a communication device, and an analysis device that uses a generative AI model. This enables elderly users to operate electronic devices quickly and accurately through real-time, intuitive voice guidance and AR displays. Specifically, the server captures images of the user's operating environment with a camera, sends the images to the server, and analyzes them with a generative AI model to provide appropriate operation guidance. Furthermore, the server converts the user's voice feedback into text for further operational assistance, making it easier to operate complex digital devices.
[0610] "Elderly users" refers to older individuals who are the primary users of the system and who have difficulty using electronic devices and software, especially those that are complex to operate.
[0611] "AR glasses" are headset-type devices that users can wear and use, equipped with augmented reality (AR) technology to display information as an overlay on their field of vision.
[0612] "Photographing means" refers to the camera built into the AR glasses, a device that has the function of capturing an image of the user interface of an electronic device within the user's field of view.
[0613] "Communication means" refers to data transmission devices such as Wi-Fi modules and 4G / 5G communication modules for transmitting captured image data to a remote server.
[0614] A "generative artificial intelligence model" is an AI model used for image recognition and natural language processing, and refers to a program that generates specific operating instructions and guides based on deep learning technology, for example.
[0615] "Analysis means" refers to software and hardware for analyzing the transmitted image data and recognizing the state of the electronic device, the position of the user interface, and the like.
[0616] "Operation guidance means" refers to a mechanism for providing users with operation procedures generated based on the analysis results through audio guidance and AR displays.
[0617] "Interactive means" refers to a device or program that has the function of receiving voice feedback from a user and generating and providing additional operation guidance based on that feedback.
[0618] "Speech recognition means" means the technology or software used to convert a user's voice feedback into text data.
[0619] "Presentation means" refers to a screen display or audio output device that visualizes and presents to the user the operation instructions generated based on the analyzed data.
[0620] This invention relates to a support system that enables elderly people to easily operate complex electronic devices. This system includes a photographing means, a communication means, an analysis means using a generative artificial intelligence model, an operation guidance means, a dialogue means, a voice recognition means, and a presentation means. The following describes in detail an embodiment of the invention.
[0621] Hardware and Software Configuration
[0622] Hardware used
[0623] 1. AR Glasses:
[0624] Built-in camera: A camera for capturing the user interface of an electronic device in the user's field of view.
[0625] Built-in microphone: A microphone for collecting user voice feedback.
[0626] Built-in speaker: A speaker for playing audio guides.
[0627] Communication module: A module that supports Wi-Fi and 4G / 5G communications.
[0628] 2. Cloud Server:
[0629] high performance computer
[0630] Generative AI models (e.g., GPT-4)
[0631] Image Analysis Software
[0632] Software used
[0633] 1. Real-time image capture software: Software that captures images from the AR glasses camera and sends them to a cloud server.
[0634] 2. Speech recognition software: Software for converting user voice feedback into text data.
[0635] 3. Generative AI model: Software that analyzes the transmitted image data and generates appropriate operating instructions.
[0636] 4. Speech synthesis software: Software for converting the generated operating instructions into audio guidance.
[0637] 5. AR display software: Software for overlaying operation guides onto the user's field of view.
[0638] System Operation
[0639] Example 1: Smartphone Wi-Fi settings
[0640] Consider a scenario where a user wears AR glasses and operates the home screen of a smartphone. When the user looks at the smartphone, the camera in the AR glasses captures the home screen, and the image data is sent to a cloud server. The cloud server analyzes this image using a generative AI model and generates an operation instruction such as "Tap the Settings app." This instruction is converted into an audio file and AR display data, which are then sent to the device. The user follows these instructions to perform the operation.
[0641] As a concrete example, the following prompt sentences could be input to a generative AI model:
[0642] "I want to set up Wi-Fi on my smartphone. Can you please tell me the detailed steps to tap the Settings app?"
[0643] Example 2: Home electronics settings
[0644] Consider a scenario where a user wears AR glasses and operates a home electronic device. For example, let's say they want to operate an air conditioner remote control. When the user picks up the remote control, the camera in the AR glasses captures the control panel, and the image is sent to a cloud server. After analysis by the generative AI model, an instruction such as "Press the power button" is generated. This instruction is provided to the user through audio guidance and AR displays.
[0645] As a concrete example, the following prompt sentences could be input to a generative AI model:
[0646] "I want to set up a home appliance. First, please tell me the steps to press the power button."
[0647] As described above, this system helps seniors easily understand and operate complex electronic devices. By providing real-time operation guidance and additional guidance based on user feedback, it is possible to reduce technological barriers and promote the use of digital devices.
[0648] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0649] Step 1:
[0650] The user puts on the AR glasses and turns on the device.
[0651] Specific operation: The user presses and holds the power button on the AR glasses to start the system, which automatically turns on the camera, microphone, speaker, etc.
[0652] Input: User action (pressing the power button)
[0653] Output: Notification of AR glasses startup status and readiness
[0654] Step 2:
[0655] Keep the device the user wants to control within sight.
[0656] Specific Action: A user brings an electronic device, such as a smartphone or consumer electronic device, into their field of view. The camera in the AR glasses captures this user interface.
[0657] Input: Placing the device in view
[0658] Output: Captured image data of the user interface
[0659] Step 3:
[0660] The device transmits the captured image data to a cloud server.
[0661] Specific operation: The device's communication means sends the captured image data to the cloud server via Wi-Fi or 4G / 5G. The user is notified of the progress during the transmission.
[0662] Input: Captured image data
[0663] Output: Notification of completion of data transmission to the cloud server
[0664] Step 4:
[0665] The server analyzes the received image data.
[0666] Specific operation: The server begins analyzing the received image data using analytical means. A generative AI model (e.g., GPT-4) is used for the analysis to recognize the device state and interface.
[0667] Input: Image data sent
[0668] Output: Parsed data (current state of the device and possible actions)
[0669] Step 5:
[0670] The server generates operation instructions using the generative AI model.
[0671] Specific actions: The generative AI model generates appropriate instructions based on the analysis results. For example, it creates specific instructions such as "Tap the Settings app."
[0672] Input: Parsed data
[0673] Output: Generated operating instructions
[0674] Step 6:
[0675] The operating procedures generated by the server are converted into audio files and AR display data.
[0676] Specific operation: The server converts the generated operation instructions into a voice file and AR display data using voice synthesis software and AR display software.
[0677] Input: Generated operation instructions
[0678] Output: Audio file and AR display data
[0679] Step 7:
[0680] The server sends the generated audio files and AR display data to the terminal.
[0681] Specific operation: The server sends the audio file and AR display data to the device using a communication method. The user is notified of the transmission status.
[0682] Input: Audio files and AR display data
[0683] Output: Notification of completion of transmission to the terminal
[0684] Step 8:
[0685] The device will guide the user on how to operate the device through voice guidance and AR displays.
[0686] Specific operation: The device plays the received audio file through the speaker and displays an overlay of operation instructions on the AR glasses display, such as "Tap the Settings app."
[0687] Input: Audio files and AR display data
[0688] Output: User guide
[0689] Step 9:
[0690] If the user has difficulty operating the device, voice feedback is provided.
[0691] Specific Action: If a user is unsure about an action, they can speak into the microphone and ask a question like, "What should I do next?" The microphone will capture their voice.
[0692] Input: User's voice feedback
[0693] Output: Captured audio data
[0694] Step 10:
[0695] The terminal converts the voice feedback into text using a voice recognition means and transmits it to a cloud server.
[0696] Specific operation: The device converts the voice data into text data using voice recognition software and sends the text data to the cloud server.
[0697] Input: Captured audio data
[0698] Output: Text data and notification of completion of transmission to the cloud server
[0699] Step 11:
[0700] The server analyzes the text data and generates additional operating instructions.
[0701] Specific operation: The server analyzes the received text data and generates additional operation instructions based on the user's question, such as "Next, select Wi-Fi."
[0702] Input: Text data
[0703] Output: Additional operating instructions
[0704] Step 12:
[0705] The server transmits the generated additional operation instructions to the terminal, and the terminal provides the user with additional guidance.
[0706] Specific operation: The server converts additional operation instructions into an audio file and AR display data and sends them to the terminal, which then provides guidance to the user based on this.
[0707] Input: Additional operating instructions
[0708] Output: Operation guide with additional audio files and AR display data
[0709] The above is the flow of processing in the program for this system and the specific operations of each.
[0710] (Application example 1)
[0711] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0712] When elderly people make electronic payments using smartphones and other digital devices, the operation is complicated and difficult for them to use. In addition, there is a high possibility that elderly people will make operational errors or face security risks, so an efficient method to support them is needed.
[0713] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0714] In this invention, the server includes an analysis means, an operation guidance means, and a dialogue means, which allows elderly users to be guided through the transaction procedure through the AR glasses when making electronic payments using their smartphones, and provides specific operation instructions according to the progress of the transaction.
[0715] "Elderly users" refer to users who are older and unfamiliar with operating digital devices.
[0716] "Wearable AR glasses" are glasses-type devices that users wear on their heads to display augmented reality.
[0717] A "user interface" refers to the part of a device that a user directly operates or inputs, such as a screen or operation panel.
[0718] "Imaging means" refers to the function of capturing images using a camera or sensor.
[0719] "Data communication means" refers to the function of transmitting captured data to other devices or servers.
[0720] A "cloud computing server" refers to a remote server used to store and analyze data over the Internet.
[0721] A "generative artificial intelligence model" is an artificial intelligence that analyzes input and data from users and generates appropriate responses and instructions.
[0722] "Analysis means" refers to a function that analyzes the captured image data and extracts information about the user's operations.
[0723] "Operation guide means" refers to a function that presents specific operation methods to the user based on the analysis results.
[0724] "Augmented reality display" refers to a technology that adds virtual information to the real field of view and displays it.
[0725] "Interaction" refers to the ability to interact with the user and provide additional guidance based on their feedback.
[0726] "Electronic payment support means" refers to a function that supports operations during electronic payment and guides transaction procedures.
[0727] "Transaction guidance means" refers to a function that provides specific operational instructions, such as scanning a QR code or tapping a payment button, depending on the progress of the transaction.
[0728] The system for realizing this invention is designed to help elderly users smoothly make electronic payments. The system is realized by the following components and operations.
[0729] System Configuration
[0730] The system mainly includes the following components:
[0731] 1. AR glasses: A device worn by elderly users that has a built-in camera, microphone, and speaker.
[0732] 2. Imaging means: The camera built into the AR glasses captures the device's user interface in real time as it comes into the user's field of view.
[0733] 3. Data communication means: Transmit the captured image data to a cloud computing server.
[0734] 4. Cloud computing server: Analyzes the received image data and generates operation guides using a generative AI model.
[0735] 5. Analysis method: Analyze the user interface in the image on the cloud server and identify the next operation to be performed.
[0736] 6. Operation guide means: Based on the analysis results, operation procedures are generated as audio guides and AR displays and provided to the user.
[0737] 7. Interaction: Receives voice feedback from the user and provides additional guidance based on that.
[0738] 8. Speech recognition means: converts the user's voice feedback into text data.
[0739] 9. Electronic payment support measures: Provide specific operating procedures for electronic payments.
[0740] 10. Transaction Guidance: Guidance on transaction operations such as scanning a QR code or tapping a payment button.
[0741] How it works
[0742] 1. Function of the imaging means: When a user wears the AR glasses and tries to operate a smartphone, the camera captures the smartphone screen and obtains image data in real time.
[0743] 2. Function of data communication means: The acquired image data is transmitted to the cloud computing server via data communication means (e.g., Wi-Fi or mobile network).
[0744] 3. Function of the analysis tool: A generative AI model on a cloud server analyzes image data, for example, identifying QR codes or specific buttons.
[0745] 4. Operation guidance function: Based on the analysis results, the cloud server generates specific operation instructions, such as "Scan the QR code" or "Tap the payment button."
[0746] 5. Audio guide and AR display: The generated operating procedure is provided to the user as an audio guide and an augmented reality display. The next operation the user should perform is presented visually and audibly.
[0747] 6. Interaction function: If the user is unsure of what to do, they can provide voice feedback (e.g., "Which button should I press next?"). This feedback is sent to the server via the microphone.
[0748] 7. Voice recognition and additional guidance: Voice feedback is converted into text data by a voice recognition tool and analyzed again by the cloud server. Additional specific operating instructions are then generated and provided again as voice guidance or AR displays.
[0749] Specific examples
[0750] Example 1: Payment processing for an electronic payment app
[0751] The user puts on the AR glasses and opens the electronic payment app. The smartphone screen is captured and analyzed by the cloud server. The server generates an instruction to "scan the QR code" and displays it as an audio guide and AR display through the AR glasses. When the user follows the instructions, the next instruction is displayed: "tap the payment button."
[0752] Prompt: A user is looking at an electronic payment app. Guide them through the next steps in the payment process.
[0753] Instructions for the generative AI model: "Scan the QR code" or "Tap the payment button"
[0754] Example 2: Feedback and additional guidance
[0755] The user is confused and asks, "Which button should I press next?" This voice input is converted into text data and sent to a cloud server. The server analyzes the feedback and generates additional instructions as the next step: "Tap the back button and reselect your payment method."
[0756] This invention enables elderly people to easily understand how to operate complex digital devices and use them safely.
[0757] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0758] Step 1:
[0759] A user puts on the AR glasses and attempts to operate a smartphone. The user launches an electronic payment app, and the camera in the AR glasses captures this screen in real time.
[0760] Input: The smartphone screen operated by the user
[0761] Output: Real-time captured image data
[0762] Specific operation: The camera in the AR glasses captures an image of the user interface and acquires it as image data.
[0763] Step 2:
[0764] The terminal transmits the acquired image data to a cloud computing server via a communication means, with the data transfer occurring via an internet connection.
[0765] Input: Image data captured in real time
[0766] Output: Image data sent to a cloud computing server
[0767] Specific operation: The communication module in the AR glasses transmits image data using Wi-Fi or a mobile network.
[0768] Step 3:
[0769] The server analyzes the received image data using an analytical method and uses a generative AI model to determine the next step to take, such as whether the user should scan a QR code or tap a payment button.
[0770] Input: Image data sent to a cloud computing server
[0771] Output: Specific operation procedures as analysis results
[0772] Specific operation: The generative artificial intelligence model analyzes the image data and generates appropriate operating instructions based on the prompt text.
[0773] Step 4:
[0774] The server generates the operating procedures as audio guidance and AR display and sends them to the terminal.
[0775] Input: Specific operation procedures as analysis results
[0776] Output: Audio guide and AR display data
[0777] Specific operation: The operation guide generation module in the server converts the instructions into audio files and AR display data.
[0778] Step 5:
[0779] The device uses the received voice guidance and AR display to guide the user on how to operate the device.
[0780] Input: Audio guide and AR display data
[0781] Output: Audio guidance and AR display to the user
[0782] What it does: Uses the speakers and display devices of the AR glasses to provide audio and visual guidance.
[0783] Step 6:
[0784] If the user is confused about an operation, they can receive voice feedback, such as asking, "Which button should I press next?" This voice feedback is collected through the microphone.
[0785] Input: User's voice feedback
[0786] Output: Collected audio data
[0787] Specific operation: The microphone inside the AR glasses captures sound and saves it as audio data.
[0788] Step 7:
[0789] The terminal converts the voice feedback into text using a voice recognition means and transmits it to a cloud server.
[0790] Input: Collected audio data
[0791] Output: Feedback data converted to text
[0792] Specific operation: Uses voice recognition technology to convert voice data into text data.
[0793] Step 8:
[0794] The server analyzes the text feedback and generates additional operating instructions, which are then provided to the user again as audio guidance and AR displays.
[0795] Input: Feedback data converted to text
[0796] Output: Audio guide and AR display data as additional operating instructions
[0797] Specific behavior: The generative AI model analyzes the feedback, generates new operating instructions, and presents them to the user again.
[0798] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0799] The present invention combines an emotion engine with a support system for elderly users to utilize generative AI. This system uses an AR glasses camera to image-recognize the user interface of the device operated by the user, and analyzes the image using a generative AI model to provide operational support using audio guidance and AR displays, as well as flexible guidance by recognizing the user's emotions. The following describes in detail the embodiments of the present invention.
[0800] System Configuration
[0801] The system mainly includes the following components:
[0802] AR glasses: Worn by elderly users, these glasses have built-in cameras, microphones, and speakers.
[0803] Camera means: Built into the AR glasses, it captures the user interface of the device within the user's field of view.
[0804] Communication means: Captured image data and audio data are sent to a cloud server.
[0805] Cloud server: Receives image data and audio data and has a generative AI model that analyzes them.
[0806] Analysis method: A generative AI model is used to analyze the user interface and audio data in the image.
[0807] Guidance means: Based on the analysis results, operating procedures are generated and presented to the user through audio guidance and AR displays.
[0808] Interaction methods: Accepts user voice feedback and provides additional guidance based on the feedback.
[0809] Speech recognition means: converts the user's voice feedback into text data.
[0810] Emotion Engine: Evaluates emotions by analyzing the user's tone of voice, facial expressions, and posture.
[0811] Program processing flow
[0812] 1. The user puts on the AR glasses
[0813] The user puts on the AR glasses and turns them on, which activates the camera and microphone and puts them into standby mode.
[0814] 2. User Actions
[0815] The user brings the device to be operated (e.g., a smartphone or a household appliance) into view, and the camera means captures the user interface of this device in real time.
[0816] 3. Data Transmission
[0817] The image and audio data captured by the device is compressed and encoded and sent to a cloud server with low latency.
[0818] 4. Cloud server analysis
[0819] The image and audio data received by the server is analyzed by an analysis means, and the user interface and audio are analyzed using a generative artificial intelligence model (generative AI).
[0820] The emotion engine analyzes voice tone, facial expressions, and posture to assess the user's emotions.
[0821] 5. Providing operation guides
[0822] Based on the analysis results and emotion evaluation, the server converts the specific next steps into text format, converts it into an audio file or AR display data, and sends it to the device.
[0823] Based on this, the device will present the user with operation instructions using voice guidance and AR display. For example, the voice guidance will say "Tap the Settings app" and the AR display will show the location of the Settings app.
[0824] 6. User Feedback
[0825] The user operates the device by following the instructions. If they are unsure of how to operate the device, they can ask questions or provide feedback by speaking into the microphone. For example, "What should I do next?"
[0826] 7. Speech Recognition and Emotion Assessment
[0827] The device captures the user's voice feedback and converts it into text data using a speech recognition tool, while the emotion engine evaluates the user's emotions.
[0828] 8. Submitting Feedback
[0829] The device sends the converted text data and emotional information to a cloud server.
[0830] 9. Response generation by cloud server
[0831] The server then re-feeds the user's questions and feedback into the generative AI model to generate an appropriate response, such as "Select your Wi-Fi settings."
[0832] Based on information from the emotion engine, the content and presentation of the response can be flexibly adjusted, for example, if the user is annoyed, a calmer tone of response can be generated.
[0833] 10. Providing a Response
[0834] The server converts the response into an audio file or AR display data and sends it to the device.
[0835] The device will again provide the response to the user through audio guidance and AR display.
[0836] Specific examples
[0837] Example 1: Smartphone Wi-Fi settings
[0838] The user puts on the AR glasses and projects the smartphone's home screen with the camera. The device sends this image to a cloud server, which then uses generative AI to analyze the image and generate the instruction, "Tap the Settings app." The AR glasses then present this instruction to the user through voice guidance and AR display. If the user is unsure of the next step, they can ask aloud, "What should I do next?" If the emotion engine detects the user's frustration along with this question, the server generates and provides more detailed instructions in a gentler tone.
[0839] Example 2: Configuring Home Appliances
[0840] The user puts on the AR glasses and brings the control panel of a household appliance into view. The device sends an image of this panel to the server, which analyzes it. For example, an instruction such as "Press the power button" is generated and presented to the user through a guide means. If the user is unsure and gives feedback such as "This doesn't work," the emotion engine will detect anxiety and the server will generate an instruction in an encouraging tone such as "Don't worry. Next, try..."
[0841] In accordance with the above aspects, the present invention realizes a system that enables elderly people to easily understand and use the operation of complex digital devices, and provides flexible support that takes emotions into consideration.
[0842] The processing flow will be explained below.
[0843] Step 1:
[0844] The user puts on the AR glasses, which activates the camera and microphone of the AR glasses and puts them into standby mode.
[0845] Step 2:
[0846] The user brings the device to be operated (e.g., a smartphone or a home appliance) into view. The device uses a camera to capture the device's user interface in real time.
[0847] Step 3:
[0848] The device compresses and encodes the captured image data and user voice data and sends them to the cloud server.
[0849] Step 4:
[0850] The server decodes the received image data, and the analysis means analyzes the user interface elements in the image using a generative artificial intelligence model (generative AI).
[0851] Step 5:
[0852] The server analyzes the received voice data and converts it into text using a voice recognition means.
[0853] Step 6:
[0854] The server uses the generation AI to generate the next steps based on the analysis results, creating specific instructions such as "tap the Settings app."
[0855] Step 7:
[0856] The server converts the generated operating procedures into audio files and AR display data and sends them to the terminal.
[0857] Step 8:
[0858] Based on the operation guide received by the device from the server, the device presents the user with operation procedures using voice guidance and AR display. For example, the voice guidance may say "Tap the Settings app" and the AR display may show the location of the Settings app.
[0859] Step 9:
[0860] The user operates the device according to the operating instructions. If the user is unsure of how to operate the device, they can ask questions into the microphone, such as "What should I do next?"
[0861] Step 10:
[0862] The device captures the user's voice feedback and converts it into text data using voice recognition technology, while the emotion engine analyzes the user's voice tone, facial expressions, and posture to make an emotional assessment.
[0863] Step 11:
[0864] The device sends the converted text data and emotional information to a cloud server.
[0865] Step 12:
[0866] The server re-inputs the user's questions and feedback into the AI model to generate an appropriate response. Based on the evaluation of the emotion engine, the model generates instructions with a tone and content that corresponds to the user's emotions. For example, if the user is showing signs of irritation, the model generates instructions in a calm tone such as "Don't worry. Next, please select your Wi-Fi settings."
[0867] Step 13:
[0868] The server converts the response into an audio file or AR display data and sends it to the device.
[0869] Step 14:
[0870] The device will again provide the response to the user through audio guidance and AR display.
[0871] Step 15:
[0872] The user resumes operation following any further instructions from the server, repeating steps 9 through 14 as necessary.
[0873] Example 2
[0874] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0875] When elderly people operate complex digital devices, they often find it difficult to understand the operating procedures, and emotions such as anxiety and frustration during operation can hinder their operation. To solve this problem, it is necessary to provide appropriate operating instructions that take into account the user's emotions.
[0876] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0877] In this invention, the server includes an interaction means for receiving voice feedback from the user and providing additional operation guidance based on the feedback, a voice recognition means for converting the user's voice feedback into text using voice recognition technology, and an emotion evaluation means for analyzing the voice tone, facial expression, and posture to evaluate the user's emotions, thereby making it possible to provide flexible and appropriate operation guidance that takes the user's emotions into consideration.
[0878] "Elderly users" refer to users who are likely to find it difficult to operate digital devices due to their age.
[0879] "AR glasses" refers to a glasses-type device that uses augmented reality technology to overlay information on the real world.
[0880] "Capture means" refers to a camera or other capture device used to capture images of objects or interfaces within the user's field of view.
[0881] "Communication means" refers to an internet connection or wireless communication means for sending and receiving image and audio data between the terminal and the cloud server.
[0882] A "generative artificial intelligence model" refers to a machine learning model designed to perform a specific task (e.g., image recognition or speech analysis).
[0883] The "analysis means" refers to a digital processor that analyzes the acquired image and audio data, recognizes the user interface, and generates operation procedures.
[0884] "Guidance means" refers to an audio output device or an augmented reality display device that notifies the user of the operating procedure generated based on the analysis results.
[0885] "Interaction means" refers to the part of the system that accepts voice or other operational input from the user and generates additional operational instructions based on that input.
[0886] "Speech recognition means" refers to technology or equipment for converting voice data obtained from a user into text data.
[0887] "Emotion assessment means" refers to technology or devices for analyzing a user's tone of voice, facial expression, posture, etc. to assess the user's emotional state.
[0888] This system aims to assist elderly users in operating complex digital devices, and uses AR glasses to capture the user interface of the target device within the user's field of view. Specifically, a camera means built into the AR glasses captures images of the user device in real time and transmits them to a cloud server via a communication means.
[0889] Hardware and software used
[0890] AR glasses: Devices that display information using augmented reality technology and have built-in cameras, microphones, and speakers.
[0891] Camera means: Capable of capturing the user interface of a device within the user's field of view.
[0892] Communication means: Includes internet connection and wireless communication means for transmitting captured image and audio data to a cloud server.
[0893] Cloud server: Receives image and audio data and analyzes them using generative artificial intelligence models.
[0894] Analysis means: A generative artificial intelligence model in the cloud server is used to analyze the user interface and audio data in the image.
[0895] Guidance means: Based on the analysis results, operating procedures are generated and presented to the user through audio guidance and augmented reality displays.
[0896] Interaction methods: Accepts user voice feedback and provides additional guidance.
[0897] Speech recognition means: Uses technology to convert the user's voice feedback into text data.
[0898] Emotion assessment means: It has technology to assess emotions by analyzing the user's tone of voice, facial expressions, and posture.
[0899] System operation explanation
[0900] When a user puts on the AR glasses and turns them on, the camera and microphone are activated and in standby mode. When the user brings the device to be operated (e.g., a smartphone, tablet, or household appliance) into view, the camera captures the device's user interface. The captured image and audio data are sent to a cloud server via a communication means.
[0901] The cloud server analyzes the received image and audio data using an analysis means and uses a generative artificial intelligence model to analyze the user interface and audio. The emotion evaluation means also analyzes the voice tone, facial expression, and posture to evaluate the user's emotions. Based on the analysis results and the emotion evaluation, specific next steps are generated, converted into audio files and augmented reality display data, and then sent to the device.
[0902] The device provides the user with operational instructions via voice guidance and augmented reality display. For example, the voice guidance might say, "Tap the Settings app," and the location of the Settings app is highlighted on the AR glasses' display. If the user is unsure of the next step, they can ask into the microphone, "What should I do next?". The emotion assessment means analyzes the user's emotions along with this question, and the cloud server generates and provides detailed, gentle-toned instructions.
[0903] Specific examples
[0904] Example 1: Smartphone Wi-Fi settings
[0905] The user puts on the AR glasses and projects the smartphone's home screen with the camera. The device sends this image to a cloud server, which then uses generative AI to analyze the image and generate the instruction, "Tap the Settings app." The AR glasses then present this instruction to the user through voice guidance and an augmented reality display. If the user is unsure of the next step, they can ask aloud, "What should I do next?" If the emotion engine detects the user's frustration along with this question, the server generates and provides more detailed instructions in a gentler tone.
[0906] Example 2: Configuring Home Appliances
[0907] The user puts on the AR glasses and brings the control panel of a household appliance into view. The device sends an image of this panel to the server, which analyzes it. For example, an instruction such as "Press the power button" is generated and presented to the user through a guide means. If the user is unsure and gives feedback such as "This doesn't work," the emotion engine will detect anxiety and the server will generate an instruction in an encouraging tone such as "Don't worry. Next, try..."
[0908] Prompt Sentence Examples
[0909] "Please provide audio guidance and AR displays to guide users through the steps they need to set up Wi-Fi on their smartphones."
[0910] The above system aims to use a generative AI model and an emotion engine to enable elderly users to easily operate complex digital devices, and to provide flexible assistance that takes into account emotions during operation.
[0911] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0912] Step 1:
[0913] The user puts on the AR glasses.
[0914] Input: The user puts on the AR glasses and turns them on.
[0915] Operation: The camera and microphone of the AR glasses are activated and enter standby mode.
[0916] Output: The camera and microphone are ready to use.
[0917] Step 2:
[0918] Keep the device the user is controlling within sight.
[0919] Input: The user brings the device they want to control (e.g., a smartphone or household appliance) into view.
[0920] Operation: A camera means captures the device's user interface in real time.
[0921] Output: Captured image and audio data.
[0922] Step 3:
[0923] Send the data to a cloud server.
[0924] Input: Captured image and audio data.
[0925] Operation: The device compresses and encodes image and audio data and sends them to the cloud server using a communication method.
[0926] Output: The cloud server receives the captured image and audio data.
[0927] Step 4:
[0928] A cloud server analyzes the image and audio data.
[0929] Input: Image data and audio data received by the cloud server.
[0930] Operation: The analysis means (generative AI model) analyzes image data and recognizes the user interface. The emotion evaluation means analyzes voice tone, facial expressions, and posture to evaluate the user's emotions.
[0931] Output: User interface analysis results and user emotion evaluation results.
[0932] Step 5:
[0933] Generate operation guide.
[0934] Input: User interface analysis results and user emotion evaluation results.
[0935] How it works: The cloud server uses the generative AI model to create specific operating procedures in text format based on the analysis results and emotional evaluation.
[0936] Output: Generated operating instructions text.
[0937] Step 6:
[0938] The device provides operation guidance.
[0939] Input: The generated operating instructions text.
[0940] How it works: The device converts the instruction text into an audio file and augmented reality display data, and presents it to the user. Specifically, the audio guide says "Tap the Settings app," and the AR display highlights the location of the Settings app.
[0941] Output: Audio guide and augmented reality display presented to the user.
[0942] Step 7:
[0943] Process user feedback.
[0944] Input: The user interacts with the device through instructions and uses voice to ask questions or provide feedback, for example, "What should I do next?"
[0945] What it does: The device captures the user's voice feedback.
[0946] Output: Captured audio feedback data.
[0947] Step 8:
[0948] Performs voice recognition and emotion assessment.
[0949] Input: Captured audio feedback data.
[0950] Operation: The device uses a voice recognition means to convert voice data into text, and an emotion assessment means analyzes the user's voice tone and facial expressions to assess their emotions.
[0951] Output: Feedback and sentiment evaluation results converted into text.
[0952] Step 9:
[0953] Send the feedback to a cloud server.
[0954] Input: Feedback and sentiment rating results converted to text.
[0955] Operation: The device sends text data and emotion information to the cloud server using a communication method.
[0956] Output: The cloud server receives the feedback data and emotion information.
[0957] Step 10:
[0958] The cloud server generates a response.
[0959] Input: Feedback data and emotional information.
[0960] How it works: The server uses a generative AI model to generate appropriate responses based on user feedback, and flexibly adjusts the content and presentation of the response based on emotional assessment, for example, generating a calmer tone of response if the user is annoyed.
[0961] Output: The generated response text.
[0962] Step 11:
[0963] The terminal provides a response.
[0964] Input: The generated response text.
[0965] How it works: The device converts the response text into an audio file and augmented reality display data, and presents them to the user. Specifically, the audio prompt says "Select Wi-Fi settings" and the corresponding part is highlighted in the AR display.
[0966] Output: Audio guide and augmented reality display presented to the user.
[0967] Through the above processing steps, this system can provide appropriate and flexible support for operating digital devices while taking into account the user's emotions.
[0968] (Application example 2)
[0969] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0970] Elderly users face challenges when operating complex digital devices or selecting products in physical stores, such as difficulty obtaining and understanding information about operation and purchasing. Additionally, elderly users often experience anxiety and frustration while operating devices, which can be difficult to deal with appropriately. This can make it difficult for elderly users to smoothly navigate digital devices and physical stores.
[0971] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0972] In this invention, the server includes a camera, a communication unit, an analysis unit using a generative AI model, an emotion evaluation unit, a guide unit, a dialogue unit, and a voice recognition unit. This allows elderly users to easily obtain and understand product information on digital devices and in physical stores using AR glasses, and also allows them to receive flexible operation guidance that takes the user's emotions into consideration.
[0973] "AR glasses" are devices worn by elderly users that can overlay digital information onto their field of vision.
[0974] "Camera means" refers to a device for capturing an image of the user interface of a device within the user's field of view.
[0975] "Communication means" is a function for transmitting captured image data to a cloud server.
[0976] A "generative AI model" is an artificial intelligence model that analyzes received image data and generates appropriate operating procedures.
[0977] "Analysis means" is a function for analyzing transmitted images using a generative AI model.
[0978] The "guide means" is a function that provides users with operating procedures generated based on the analysis results through audio guidance and AR displays.
[0979] The "emotion evaluation means" is a function for analyzing the user's voice and facial expressions and evaluating the user's emotional state.
[0980] "Interactive means" is a function that accepts voice feedback from the user and provides additional operational guidance and emotion-sensitive guidance based on that feedback.
[0981] "Voice recognition means" is a function for converting the user's voice feedback into text data.
[0982] This invention is a system that makes it easier for elderly users to select products using complex digital devices or in physical stores. The system includes AR glasses worn by the elderly user, a camera means, a communication means, an analysis means having a generative AI model, an emotion evaluation means, a guidance means, a dialogue means, and a voice recognition means.
[0983] System Configuration
[0984] 1. AR Glasses:
[0985] This device is worn by the user and has a built-in camera, microphone, and speaker, allowing digital information to be displayed overlaid on the user's field of vision when operating a digital device or selecting products in a physical store.
[0986] 2. Camera means:
[0987] It is built into AR glasses and captures images of the device's user interface and products in physical stores within the user's field of view.
[0988] 3. Means of communication:
[0989] The captured image and audio data are sent to a cloud server using a secure, low-latency protocol.
[0990] 4. Analytical means with generative AI models:
[0991] The image data received by the server is analyzed and appropriate operating procedures and product information are generated. The generative AI model learns from a large data set and provides highly accurate analysis results.
[0992] 5. Emotional assessment measures:
[0993] It analyzes the user's tone of voice, facial expressions, and posture to assess the user's emotional state, allowing it to process user feedback appropriately and provide emotionally sensitive assistance.
[0994] 6. Guide means:
[0995] Based on the analysis results, the system generates specific next steps and product information, and provides these instructions to the user in the form of audio guidance or AR displays.
[0996] 7. Means of interaction:
[0997] If the user is unsure of a procedure or needs additional information, the system accepts voice feedback and provides additional operating instructions or emotionally sensitive guidance based on the feedback.
[0998] 8. Voice Recognition Methods:
[0999] The user's voice feedback is converted into text data, which is further analyzed by generative AI models and sentiment assessment methods.
[1000] Example
[1001] Example 1: Smartphone settings
[1002] The user puts on the AR glasses and projects the smartphone's home screen onto the camera. The cloud server analyzes this image using a generative AI model and provides the instruction, "Tap the Settings app." The AR glasses present this instruction to the user through voice guidance and AR displays. If the user is unsure of the next step, they can ask aloud, "What should I do next?" If the emotion engine detects the user's frustration along with this question, the server generates and provides more detailed instructions in a gentler tone.
[1003] Example 2: Product selection in a store
[1004] A user enters a supermarket and wants to look at a particular shelf. The AR glasses send an image of the shelf to a cloud server. The cloud server analyzes the image using a generative AI model and provides a voice prompt saying, "This product is rich in probiotics." If the user's emotions are unstable, the emotion assessment mechanism detects this and provides instructions in an encouraging tone, such as, "Don't worry, check out this product next."
[1005] Prompt example
[1006] "What should I do next?"
[1007] "Is this the right button to press?"
[1008] As described above, this system allows elderly users to easily operate digital devices and select products in physical stores. Furthermore, the emotional support allows them to use the system with peace of mind.
[1009] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1010] Step 1:
[1011] A user wears AR glasses and places a digital device or product in their field of view. The camera built into the AR glasses captures an image of the user interface or product. This image data becomes the input for the system. The camera means performs real-time image capture.
[1012] Step 2:
[1013] The device sends image data captured by the camera to the cloud server via a communication means. At this time, the image data is compressed, encoded, and sent with low latency. Data transmission is performed safely and quickly using a communication protocol.
[1014] Step 3:
[1015] The server analyzes the received image data using a generative AI model. This analysis method identifies the user interface and product attributes in the image and generates operation procedures and product information as the analysis results. The input is image data, and the output is the analysis results.
[1016] Step 4:
[1017] Based on the analysis results, the server generates the next operation procedure and product information, which the guide means provides to the terminal in the form of audio guide and AR display. The guide means converts the generated text into an audio file and AR display data.
[1018] Step 5:
[1019] The user operates the device and checks the product according to the presented operating instructions and product information. If the user is unsure of the next step, voice feedback is provided through the device's microphone. A prompt such as "What should I do next?" is input.
[1020] Step 6:
[1021] The device captures the user's voice feedback and converts it into text data using a voice recognition means. At the same time, the emotion assessment means analyzes the user's voice tone and facial expressions to evaluate the user's emotional state. The input is voice feedback, and the output is text data and emotion information.
[1022] Step 7:
[1023] The text data and emotional information converted by the speech recognition means are sent to a cloud server. The server then inputs this data back into the generative AI model to generate an appropriate response. The content and presentation of the response are flexibly adjusted based on the analysis results of the emotional evaluation means.
[1024] Step 8:
[1025] The server converts the generated response into an audio file and AR display data and sends them to the terminal. Based on this, the guide means provides the user with a response again in the form of audio guidance and AR display. For example, an instruction such as "Please select Wi-Fi settings" is generated.
[1026] This series of steps allows elderly users to smoothly operate digital devices and select products in physical stores.
[1027] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1028] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1029] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1030] [Third embodiment]
[1031] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1032] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1034] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1035] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1036] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1038] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1039] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1041] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1042] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1043] The present invention relates to a support system for elderly people to utilize generative AI and IT tools. This system uses an AR glasses camera to image-recognize the user interface of the device operated by the user, and analyzes the image using a generative AI model to provide operational support using audio guidance and AR displays. The following describes in detail the embodiments of the present invention.
[1044] System Configuration
[1045] The system mainly includes the following components:
[1046] AR glasses: Worn by elderly users, these glasses have built-in cameras, microphones, and speakers.
[1047] Camera means: Built into the AR glasses, it captures the user interface of the device within the user's field of view.
[1048] Communication method: Captured image data is sent to a cloud server.
[1049] Cloud server: Receives image data and has a generative AI model that performs analysis.
[1050] Analysis method: Analyze the user interface in the image using a generative AI model.
[1051] Guidance means: Operating procedures are generated based on the analysis results and presented to the user through audio guidance and AR displays.
[1052] Interaction methods: Receives user voice feedback and generates additional operation instructions based on the feedback.
[1053] Speech recognition means: converts the user's voice feedback into text data.
[1054] Program processing flow
[1055] 1. The user puts on the AR glasses
[1056] The user puts on the AR glasses and turns on the device, which activates the camera, microphone, and other components.
[1057] 2. User Operation
[1058] The user places the device (e.g., smartphone, tablet, or household appliance) they want to operate within their field of view, and the camera captures the user interface of this device in real time.
[1059] 3. Data Transmission
[1060] The terminal transmits the captured image data to a cloud server using a communication means.
[1061] 4. Cloud server analysis
[1062] The server analyzes the received image data using an analysis means and generates device operation instructions using a generative AI model, such as "Tap the Settings app."
[1063] 5. Provision of guides
[1064] The server converts the generated operating procedures into audio files and AR display data and sends them to the terminal.
[1065] Based on this, the device will use voice guidance and AR displays to guide the user on how to operate the device.
[1066] 6. User Feedback
[1067] If the user is confused about an operation, the device will provide voice feedback through the microphone, for example, asking, "What should I do next?"
[1068] 7. Voice Recognition and Further Guidance
[1069] The terminal converts the voice feedback into text using a voice recognition means and transmits it to a cloud server.
[1070] The server analyzes this feedback, generates additional operating instructions, and again uses the guide means to provide operating assistance to the user.
[1071] Specific examples
[1072] Example 1: Smartphone Wi-Fi settings
[1073] The user puts on the AR glasses and projects the smartphone's home screen with the camera. The device sends this image to a cloud server, which then uses generative AI to analyze the image and generate the instruction, "Tap the Settings app." The AR glasses then present this instruction to the user through audio guidance and AR displays. If the user is unsure of the next step, they can ask, "What should I do next?" and the server will continue to generate and provide additional instructions.
[1074] Example 2: Configuring Home Appliances
[1075] The user puts on the AR glasses and places the control panel of a household appliance in their field of view. The device sends an image of this panel to the server, which analyzes it. For example, an instruction such as "Press the power button" is generated and presented to the user through the guidance means. If the user is still unsure about how to operate the device, assistance continues through similar voice feedback and analysis.
[1076] In accordance with the above aspects, the present invention realizes a system that enables elderly people to easily understand and use the operation of complex digital devices.
[1077] The processing flow will be explained below.
[1078] Step 1:
[1079] The user puts on the AR glasses and turns on the device, which activates the camera and microphone and puts them into standby mode.
[1080] Step 2:
[1081] The user brings the device to be operated (e.g., a smartphone or a household appliance) into view, and the camera means captures the user interface of the device in real time.
[1082] Step 3:
[1083] The device compresses and encodes the captured images and sends them to a cloud server with low latency.
[1084] Step 4:
[1085] The server decodes the image data sent and inputs it into a generative artificial intelligence model (generative AI) to begin analysis.
[1086] Step 5:
[1087] A generative AI built into the server analyzes UI elements in the image and identifies the device's current state (e.g., "home screen" or "settings screen").
[1088] Step 6:
[1089] Based on the analysis results, the server generates specific instructions in text format for the next steps, such as "Tap the settings icon."
[1090] Step 7:
[1091] The operating procedures generated by the server are converted into audio files and AR display data and sent to the terminal.
[1092] Step 8:
[1093] Based on the operation guide received by the device from the server, the device presents the user with operation procedures using voice guidance and AR display. For example, the voice guidance may say "Tap the Settings app" and the AR display may show the location of the Settings app.
[1094] Step 9:
[1095] The user operates the device according to the operating instructions. If the user is unsure of an operation, they can ask questions or provide feedback by speaking into the microphone. For example, "What should I do next?"
[1096] Step 10:
[1097] The device captures the user's voice and converts it into text data using voice recognition technology.
[1098] Step 11:
[1099] The terminal transmits the converted text data to the server.
[1100] Step 12:
[1101] The server then re-feeds the user's questions and feedback into the generative AI model to generate an appropriate response, such as "Select your Wi-Fi settings."
[1102] Step 13:
[1103] The server converts the response into an audio file or AR display data and sends it to the device.
[1104] Step 14:
[1105] The device will again provide the response to the user through audio guidance and AR display.
[1106] Step 15:
[1107] The user resumes operation according to further instructions from the server, repeating steps 9 through 14 as necessary.
[1108] Example 1
[1109] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1110] Currently, the difficulty that older people experience when operating complex electronic devices and software is a major problem. Understanding new operation methods and interfaces is a challenge, especially when introducing new technologies and tools, creating technical barriers. To address this issue, assistance systems that can be operated intuitively by older people are needed, and the importance of real-time operation guidance and feedback systems is particularly increasing.
[1111] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1112] In this invention, the server includes a camera, a communication device, and an analysis device that uses a generative AI model. This enables elderly users to operate electronic devices quickly and accurately through real-time, intuitive voice guidance and AR displays. Specifically, the server captures images of the user's operating environment with a camera, sends the images to the server, and analyzes them with a generative AI model to provide appropriate operation guidance. Furthermore, the server converts the user's voice feedback into text for further operational assistance, making it easier to operate complex digital devices.
[1113] "Elderly users" refers to older individuals who are the primary users of the system and who have difficulty using electronic devices and software, especially those that are complex to operate.
[1114] "AR glasses" are headset-type devices that users can wear and use, equipped with augmented reality (AR) technology to display information as an overlay on their field of vision.
[1115] "Photographing means" refers to the camera built into the AR glasses, a device that has the function of capturing an image of the user interface of an electronic device within the user's field of view.
[1116] "Communication means" refers to data transmission devices such as Wi-Fi modules and 4G / 5G communication modules for transmitting captured image data to a remote server.
[1117] A "generative artificial intelligence model" is an AI model used for image recognition and natural language processing, and refers to a program that generates specific operating instructions and guides based on deep learning technology, for example.
[1118] "Analysis means" refers to software and hardware for analyzing the transmitted image data and recognizing the state of the electronic device, the position of the user interface, and the like.
[1119] "Operation guidance means" refers to a mechanism for providing users with operation procedures generated based on the analysis results through audio guidance and AR displays.
[1120] "Interactive means" refers to a device or program that has the function of receiving voice feedback from a user and generating and providing additional operation guidance based on that feedback.
[1121] "Speech recognition means" means the technology or software used to convert a user's voice feedback into text data.
[1122] "Presentation means" refers to a screen display or audio output device that visualizes and presents to the user the operation instructions generated based on the analyzed data.
[1123] This invention relates to a support system that enables elderly people to easily operate complex electronic devices. This system includes a photographing means, a communication means, an analysis means using a generative artificial intelligence model, an operation guidance means, a dialogue means, a voice recognition means, and a presentation means. The following describes in detail an embodiment of the invention.
[1124] Hardware and Software Configuration
[1125] Hardware used
[1126] 1. AR Glasses:
[1127] Built-in camera: A camera for capturing the user interface of an electronic device in the user's field of view.
[1128] Built-in microphone: A microphone for collecting user voice feedback.
[1129] Built-in speaker: A speaker for playing audio guides.
[1130] Communication module: A module that supports Wi-Fi and 4G / 5G communications.
[1131] 2. Cloud Server:
[1132] high performance computer
[1133] Generative AI models (e.g., GPT-4)
[1134] Image Analysis Software
[1135] Software used
[1136] 1. Real-time image capture software: Software that captures images from the AR glasses camera and sends them to a cloud server.
[1137] 2. Speech recognition software: Software for converting user voice feedback into text data.
[1138] 3. Generative AI model: Software that analyzes the transmitted image data and generates appropriate operating instructions.
[1139] 4. Speech synthesis software: Software for converting the generated operating instructions into audio guidance.
[1140] 5. AR display software: Software for overlaying operation guides onto the user's field of view.
[1141] System Operation
[1142] Example 1: Smartphone Wi-Fi settings
[1143] Consider a scenario where a user wears AR glasses and operates the home screen of a smartphone. When the user looks at the smartphone, the camera in the AR glasses captures the home screen, and the image data is sent to a cloud server. The cloud server analyzes this image using a generative AI model and generates an operation instruction such as "Tap the Settings app." This instruction is converted into an audio file and AR display data, which are then sent to the device. The user follows these instructions to perform the operation.
[1144] As a concrete example, the following prompt sentences could be input to a generative AI model:
[1145] "I want to set up Wi-Fi on my smartphone. Can you please tell me the detailed steps to tap the Settings app?"
[1146] Example 2: Home electronics settings
[1147] Consider a scenario where a user wears AR glasses and operates a home electronic device. For example, let's say they want to operate an air conditioner remote control. When the user picks up the remote control, the camera in the AR glasses captures the control panel, and the image is sent to a cloud server. After analysis by the generative AI model, an instruction such as "Press the power button" is generated. This instruction is provided to the user through audio guidance and AR displays.
[1148] As a concrete example, the following prompt sentences could be input to a generative AI model:
[1149] "I want to set up a home appliance. First, please tell me the steps to press the power button."
[1150] As described above, this system helps seniors easily understand and operate complex electronic devices. By providing real-time operation guidance and additional guidance based on user feedback, it is possible to reduce technological barriers and promote the use of digital devices.
[1151] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1152] Step 1:
[1153] The user puts on the AR glasses and turns on the device.
[1154] Specific operation: The user presses and holds the power button on the AR glasses to start the system, which automatically turns on the camera, microphone, speaker, etc.
[1155] Input: User action (pressing the power button)
[1156] Output: Notification of AR glasses startup status and readiness
[1157] Step 2:
[1158] Keep the device the user wants to control within sight.
[1159] Specific Action: A user brings an electronic device, such as a smartphone or consumer electronic device, into their field of view. The camera in the AR glasses captures this user interface.
[1160] Input: Placing the device in view
[1161] Output: Captured image data of the user interface
[1162] Step 3:
[1163] The device transmits the captured image data to a cloud server.
[1164] Specific operation: The device's communication means sends the captured image data to the cloud server via Wi-Fi or 4G / 5G. The user is notified of the progress during the transmission.
[1165] Input: Captured image data
[1166] Output: Notification of completion of data transmission to the cloud server
[1167] Step 4:
[1168] The server analyzes the received image data.
[1169] Specific operation: The server begins analyzing the received image data using analytical means. A generative AI model (e.g., GPT-4) is used for the analysis to recognize the device state and interface.
[1170] Input: Image data sent
[1171] Output: Parsed data (current state of the device and possible actions)
[1172] Step 5:
[1173] The server generates operation instructions using the generative AI model.
[1174] Specific actions: The generative AI model generates appropriate instructions based on the analysis results. For example, it creates specific instructions such as "Tap the Settings app."
[1175] Input: Parsed data
[1176] Output: Generated operating instructions
[1177] Step 6:
[1178] The operating procedures generated by the server are converted into audio files and AR display data.
[1179] Specific operation: The server converts the generated operation instructions into a voice file and AR display data using voice synthesis software and AR display software.
[1180] Input: Generated operation instructions
[1181] Output: Audio file and AR display data
[1182] Step 7:
[1183] The server sends the generated audio files and AR display data to the terminal.
[1184] Specific operation: The server sends the audio file and AR display data to the device using a communication method. The user is notified of the transmission status.
[1185] Input: Audio files and AR display data
[1186] Output: Notification of completion of transmission to the terminal
[1187] Step 8:
[1188] The device will guide the user on how to operate the device through voice guidance and AR displays.
[1189] Specific operation: The device plays the received audio file through the speaker and displays an overlay of operation instructions on the AR glasses display, such as "Tap the Settings app."
[1190] Input: Audio files and AR display data
[1191] Output: User guide
[1192] Step 9:
[1193] If the user has difficulty operating the device, voice feedback is provided.
[1194] Specific Action: If a user is unsure about an action, they can speak into the microphone and ask a question like, "What should I do next?" The microphone will capture their voice.
[1195] Input: User's voice feedback
[1196] Output: Captured audio data
[1197] Step 10:
[1198] The terminal converts the voice feedback into text using a voice recognition means and transmits it to a cloud server.
[1199] Specific operation: The device converts the voice data into text data using voice recognition software and sends the text data to the cloud server.
[1200] Input: Captured audio data
[1201] Output: Text data and notification of completion of transmission to the cloud server
[1202] Step 11:
[1203] The server analyzes the text data and generates additional operating instructions.
[1204] Specific operation: The server analyzes the received text data and generates additional operation instructions based on the user's question, such as "Next, select Wi-Fi."
[1205] Input: Text data
[1206] Output: Additional operating instructions
[1207] Step 12:
[1208] The server transmits the generated additional operation instructions to the terminal, and the terminal provides the user with additional guidance.
[1209] Specific operation: The server converts additional operation instructions into an audio file and AR display data and sends them to the terminal, which then provides guidance to the user based on this.
[1210] Input: Additional operating instructions
[1211] Output: Operation guide with additional audio files and AR display data
[1212] The above is the flow of processing in the program for this system and the specific operations of each.
[1213] (Application example 1)
[1214] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1215] When elderly people make electronic payments using smartphones and other digital devices, the operation is complicated and difficult for them to use. In addition, there is a high possibility that elderly people will make operational errors or face security risks, so an efficient method to support them is needed.
[1216] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1217] In this invention, the server includes an analysis means, an operation guidance means, and a dialogue means, which allows elderly users to be guided through the transaction procedure through the AR glasses when making electronic payments using their smartphones, and provides specific operation instructions according to the progress of the transaction.
[1218] "Elderly users" refer to users who are older and unfamiliar with operating digital devices.
[1219] "Wearable AR glasses" are glasses-type devices that users wear on their heads to display augmented reality.
[1220] A "user interface" refers to the part of a device that a user directly operates or inputs, such as a screen or operation panel.
[1221] "Imaging means" refers to the function of capturing images using a camera or sensor.
[1222] "Data communication means" refers to the function of transmitting captured data to other devices or servers.
[1223] A "cloud computing server" refers to a remote server used to store and analyze data over the Internet.
[1224] A "generative artificial intelligence model" is an artificial intelligence that analyzes input and data from users and generates appropriate responses and instructions.
[1225] "Analysis means" refers to a function that analyzes the captured image data and extracts information about the user's operations.
[1226] "Operation guide means" refers to a function that presents specific operation methods to the user based on the analysis results.
[1227] "Augmented reality display" refers to a technology that adds virtual information to the real field of view and displays it.
[1228] "Interaction" refers to the ability to interact with the user and provide additional guidance based on their feedback.
[1229] "Electronic payment support means" refers to a function that supports operations during electronic payment and guides transaction procedures.
[1230] "Transaction guidance means" refers to a function that provides specific operational instructions, such as scanning a QR code or tapping a payment button, depending on the progress of the transaction.
[1231] The system for realizing this invention is designed to help elderly users smoothly make electronic payments. The system is realized by the following components and operations.
[1232] System Configuration
[1233] The system mainly includes the following components:
[1234] 1. AR glasses: A device worn by elderly users that has a built-in camera, microphone, and speaker.
[1235] 2. Imaging means: The camera built into the AR glasses captures the device's user interface in real time as it comes into the user's field of view.
[1236] 3. Data communication means: Transmit the captured image data to a cloud computing server.
[1237] 4. Cloud computing server: Analyzes the received image data and generates operation guides using a generative AI model.
[1238] 5. Analysis method: Analyze the user interface in the image on the cloud server and identify the next operation to be performed.
[1239] 6. Operation guide means: Based on the analysis results, operation procedures are generated as audio guides and AR displays and provided to the user.
[1240] 7. Interaction: Receives voice feedback from the user and provides additional guidance based on that.
[1241] 8. Speech recognition means: converts the user's voice feedback into text data.
[1242] 9. Electronic payment support measures: Provide specific operating procedures for electronic payments.
[1243] 10. Transaction Guidance: Guidance on transaction operations such as scanning a QR code or tapping a payment button.
[1244] How it works
[1245] 1. Function of the imaging means: When a user wears the AR glasses and tries to operate a smartphone, the camera captures the smartphone screen and obtains image data in real time.
[1246] 2. Function of data communication means: The acquired image data is transmitted to the cloud computing server via data communication means (e.g., Wi-Fi or mobile network).
[1247] 3. Function of the analysis tool: A generative AI model on a cloud server analyzes image data, for example, identifying QR codes or specific buttons.
[1248] 4. Operation guidance function: Based on the analysis results, the cloud server generates specific operation instructions, such as "Scan the QR code" or "Tap the payment button."
[1249] 5. Audio guide and AR display: The generated operating procedure is provided to the user as an audio guide and an augmented reality display. The next operation the user should perform is presented visually and audibly.
[1250] 6. Interaction function: If the user is unsure of what to do, they can provide voice feedback (e.g., "Which button should I press next?"). This feedback is sent to the server via the microphone.
[1251] 7. Voice recognition and additional guidance: Voice feedback is converted into text data by a voice recognition tool and analyzed again by the cloud server. Additional specific operating instructions are then generated and provided again as voice guidance or AR displays.
[1252] Specific examples
[1253] Example 1: Payment processing for an electronic payment app
[1254] The user puts on the AR glasses and opens the electronic payment app. The smartphone screen is captured and analyzed by the cloud server. The server generates an instruction to "scan the QR code" and displays it as an audio guide and AR display through the AR glasses. When the user follows the instructions, the next instruction is displayed: "tap the payment button."
[1255] Prompt: A user is looking at an electronic payment app. Guide them through the next steps in the payment process.
[1256] Instructions for the generative AI model: "Scan the QR code" or "Tap the payment button"
[1257] Example 2: Feedback and additional guidance
[1258] The user is confused and asks, "Which button should I press next?" This voice input is converted into text data and sent to a cloud server. The server analyzes the feedback and generates additional instructions as the next step: "Tap the back button and reselect your payment method."
[1259] This invention enables elderly people to easily understand how to operate complex digital devices and use them safely.
[1260] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1261] Step 1:
[1262] A user puts on the AR glasses and attempts to operate a smartphone. The user launches an electronic payment app, and the camera in the AR glasses captures this screen in real time.
[1263] Input: The smartphone screen operated by the user
[1264] Output: Real-time captured image data
[1265] Specific operation: The camera in the AR glasses captures an image of the user interface and acquires it as image data.
[1266] Step 2:
[1267] The terminal transmits the acquired image data to a cloud computing server via a communication means, with the data transfer occurring via an internet connection.
[1268] Input: Image data captured in real time
[1269] Output: Image data sent to a cloud computing server
[1270] Specific operation: The communication module in the AR glasses transmits image data using Wi-Fi or a mobile network.
[1271] Step 3:
[1272] The server analyzes the received image data using an analytical method and uses a generative AI model to determine the next step to take, such as whether the user should scan a QR code or tap a payment button.
[1273] Input: Image data sent to a cloud computing server
[1274] Output: Specific operation procedures as analysis results
[1275] Specific operation: The generative artificial intelligence model analyzes the image data and generates appropriate operating instructions based on the prompt text.
[1276] Step 4:
[1277] The server generates the operating procedures as audio guidance and AR display and sends them to the terminal.
[1278] Input: Specific operation procedures as analysis results
[1279] Output: Audio guide and AR display data
[1280] Specific operation: The operation guide generation module in the server converts the instructions into audio files and AR display data.
[1281] Step 5:
[1282] The device uses the received voice guidance and AR display to guide the user on how to operate the device.
[1283] Input: Audio guide and AR display data
[1284] Output: Audio guidance and AR display to the user
[1285] What it does: Uses the speakers and display devices of the AR glasses to provide audio and visual guidance.
[1286] Step 6:
[1287] If the user is confused about an operation, they can receive voice feedback, such as asking, "Which button should I press next?" This voice feedback is collected through the microphone.
[1288] Input: User's voice feedback
[1289] Output: Collected audio data
[1290] Specific operation: The microphone inside the AR glasses captures sound and saves it as audio data.
[1291] Step 7:
[1292] The terminal converts the voice feedback into text using a voice recognition means and transmits it to a cloud server.
[1293] Input: Collected audio data
[1294] Output: Feedback data converted to text
[1295] Specific operation: Uses voice recognition technology to convert voice data into text data.
[1296] Step 8:
[1297] The server analyzes the text feedback and generates additional operating instructions, which are then provided to the user again as audio guidance and AR displays.
[1298] Input: Feedback data converted to text
[1299] Output: Audio guide and AR display data as additional operating instructions
[1300] Specific behavior: The generative AI model analyzes the feedback, generates new operating instructions, and presents them to the user again.
[1301] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1302] The present invention combines an emotion engine with a support system for elderly users to utilize generative AI. This system uses an AR glasses camera to image-recognize the user interface of the device operated by the user, and analyzes the image using a generative AI model to provide operational support using audio guidance and AR displays, as well as flexible guidance by recognizing the user's emotions. The following describes in detail the embodiments of the present invention.
[1303] System Configuration
[1304] The system mainly includes the following components:
[1305] AR glasses: Worn by elderly users, these glasses have built-in cameras, microphones, and speakers.
[1306] Camera means: Built into the AR glasses, it captures the user interface of the device within the user's field of view.
[1307] Communication means: Captured image data and audio data are sent to a cloud server.
[1308] Cloud server: Receives image data and audio data and has a generative AI model that analyzes them.
[1309] Analysis method: A generative AI model is used to analyze the user interface and audio data in the image.
[1310] Guidance means: Based on the analysis results, operating procedures are generated and presented to the user through audio guidance and AR displays.
[1311] Interaction methods: Accepts user voice feedback and provides additional guidance based on the feedback.
[1312] Speech recognition means: converts the user's voice feedback into text data.
[1313] Emotion Engine: Evaluates emotions by analyzing the user's tone of voice, facial expressions, and posture.
[1314] Program processing flow
[1315] 1. The user puts on the AR glasses
[1316] The user puts on the AR glasses and turns them on, which activates the camera and microphone and puts them into standby mode.
[1317] 2. User Actions
[1318] The user brings the device to be operated (e.g., a smartphone or a household appliance) into view, and the camera means captures the user interface of this device in real time.
[1319] 3. Data Transmission
[1320] The image and audio data captured by the device is compressed and encoded and sent to a cloud server with low latency.
[1321] 4. Cloud server analysis
[1322] The image and audio data received by the server is analyzed by an analysis means, and the user interface and audio are analyzed using a generative artificial intelligence model (generative AI).
[1323] The emotion engine analyzes voice tone, facial expressions, and posture to assess the user's emotions.
[1324] 5. Providing operation guides
[1325] Based on the analysis results and emotion evaluation, the server converts the specific next steps into text format, converts it into an audio file or AR display data, and sends it to the device.
[1326] Based on this, the device will present the user with operation instructions using voice guidance and AR display. For example, the voice guidance will say "Tap the Settings app" and the AR display will show the location of the Settings app.
[1327] 6. User Feedback
[1328] The user operates the device by following the instructions. If they are unsure of how to operate the device, they can ask questions or provide feedback by speaking into the microphone. For example, "What should I do next?"
[1329] 7. Speech Recognition and Emotion Assessment
[1330] The device captures the user's voice feedback and converts it into text data using a speech recognition tool, while the emotion engine evaluates the user's emotions.
[1331] 8. Submitting Feedback
[1332] The device sends the converted text data and emotional information to a cloud server.
[1333] 9. Response generation by cloud server
[1334] The server then re-feeds the user's questions and feedback into the generative AI model to generate an appropriate response, such as "Select your Wi-Fi settings."
[1335] Based on information from the emotion engine, the content and presentation of the response can be flexibly adjusted, for example, if the user is annoyed, a calmer tone of response can be generated.
[1336] 10. Providing a Response
[1337] The server converts the response into an audio file or AR display data and sends it to the device.
[1338] The device will again provide the response to the user through audio guidance and AR display.
[1339] Specific examples
[1340] Example 1: Smartphone Wi-Fi settings
[1341] The user puts on the AR glasses and projects the smartphone's home screen with the camera. The device sends this image to a cloud server, which then uses generative AI to analyze the image and generate the instruction, "Tap the Settings app." The AR glasses then present this instruction to the user through voice guidance and AR display. If the user is unsure of the next step, they can ask aloud, "What should I do next?" If the emotion engine detects the user's frustration along with this question, the server generates and provides more detailed instructions in a gentler tone.
[1342] Example 2: Configuring Home Appliances
[1343] The user puts on the AR glasses and brings the control panel of a household appliance into view. The device sends an image of this panel to the server, which analyzes it. For example, an instruction such as "Press the power button" is generated and presented to the user through a guide means. If the user is unsure and gives feedback such as "This doesn't work," the emotion engine will detect anxiety and the server will generate an instruction in an encouraging tone such as "Don't worry. Next, try..."
[1344] In accordance with the above aspects, the present invention realizes a system that enables elderly people to easily understand and use the operation of complex digital devices, and provides flexible support that takes emotions into consideration.
[1345] The processing flow will be explained below.
[1346] Step 1:
[1347] The user puts on the AR glasses, which activates the camera and microphone of the AR glasses and puts them into standby mode.
[1348] Step 2:
[1349] The user brings the device to be operated (e.g., a smartphone or a home appliance) into view. The device uses a camera to capture the device's user interface in real time.
[1350] Step 3:
[1351] The device compresses and encodes the captured image data and user voice data and sends them to the cloud server.
[1352] Step 4:
[1353] The server decodes the received image data, and the analysis means analyzes the user interface elements in the image using a generative artificial intelligence model (generative AI).
[1354] Step 5:
[1355] The server analyzes the received voice data and converts it into text using a voice recognition means.
[1356] Step 6:
[1357] The server uses the generation AI to generate the next steps based on the analysis results, creating specific instructions such as "tap the Settings app."
[1358] Step 7:
[1359] The server converts the generated operating procedures into audio files and AR display data and sends them to the terminal.
[1360] Step 8:
[1361] Based on the operation guide received by the device from the server, the device presents the user with operation procedures using voice guidance and AR display. For example, the voice guidance may say "Tap the Settings app" and the AR display may show the location of the Settings app.
[1362] Step 9:
[1363] The user operates the device according to the operating instructions. If the user is unsure of how to operate the device, they can ask questions into the microphone, such as "What should I do next?"
[1364] Step 10:
[1365] The device captures the user's voice feedback and converts it into text data using voice recognition technology, while the emotion engine analyzes the user's voice tone, facial expressions, and posture to make an emotional assessment.
[1366] Step 11:
[1367] The device sends the converted text data and emotional information to a cloud server.
[1368] Step 12:
[1369] The server re-inputs the user's questions and feedback into the AI model to generate an appropriate response. Based on the evaluation of the emotion engine, the model generates instructions with a tone and content that corresponds to the user's emotions. For example, if the user is showing signs of irritation, the model generates instructions in a calm tone such as "Don't worry. Next, please select your Wi-Fi settings."
[1370] Step 13:
[1371] The server converts the response into an audio file or AR display data and sends it to the device.
[1372] Step 14:
[1373] The device will again provide the response to the user through audio guidance and AR display.
[1374] Step 15:
[1375] The user resumes operation following any further instructions from the server, repeating steps 9 through 14 as necessary.
[1376] Example 2
[1377] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1378] When elderly people operate complex digital devices, they often find it difficult to understand the operating procedures, and emotions such as anxiety and frustration during operation can hinder their operation. To solve this problem, it is necessary to provide appropriate operating instructions that take into account the user's emotions.
[1379] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1380] In this invention, the server includes an interaction means for receiving voice feedback from the user and providing additional operation guidance based on the feedback, a voice recognition means for converting the user's voice feedback into text using voice recognition technology, and an emotion evaluation means for analyzing the voice tone, facial expression, and posture to evaluate the user's emotions, thereby making it possible to provide flexible and appropriate operation guidance that takes the user's emotions into consideration.
[1381] "Elderly users" refer to users who are likely to find it difficult to operate digital devices due to their age.
[1382] "AR glasses" refers to a glasses-type device that uses augmented reality technology to overlay information on the real world.
[1383] "Capture means" refers to a camera or other capture device used to capture images of objects or interfaces within the user's field of view.
[1384] "Communication means" refers to an internet connection or wireless communication means for sending and receiving image and audio data between the terminal and the cloud server.
[1385] A "generative artificial intelligence model" refers to a machine learning model designed to perform a specific task (e.g., image recognition or speech analysis).
[1386] The "analysis means" refers to a digital processor that analyzes the acquired image and audio data, recognizes the user interface, and generates operation procedures.
[1387] "Guidance means" refers to an audio output device or an augmented reality display device that notifies the user of the operating procedure generated based on the analysis results.
[1388] "Interaction means" refers to the part of the system that accepts voice or other operational input from the user and generates additional operational instructions based on that input.
[1389] "Speech recognition means" refers to technology or equipment for converting voice data obtained from a user into text data.
[1390] "Emotion assessment means" refers to technology or devices for analyzing a user's tone of voice, facial expression, posture, etc. to assess the user's emotional state.
[1391] This system aims to assist elderly users in operating complex digital devices, and uses AR glasses to capture the user interface of the target device within the user's field of view. Specifically, a camera means built into the AR glasses captures images of the user device in real time and transmits them to a cloud server via a communication means.
[1392] Hardware and software used
[1393] AR glasses: Devices that display information using augmented reality technology and have built-in cameras, microphones, and speakers.
[1394] Camera means: Capable of capturing the user interface of a device within the user's field of view.
[1395] Communication means: Includes internet connection and wireless communication means for transmitting captured image and audio data to a cloud server.
[1396] Cloud server: Receives image and audio data and analyzes them using generative artificial intelligence models.
[1397] Analysis means: A generative artificial intelligence model in the cloud server is used to analyze the user interface and audio data in the image.
[1398] Guidance means: Based on the analysis results, operating procedures are generated and presented to the user through audio guidance and augmented reality displays.
[1399] Interaction methods: Accepts user voice feedback and provides additional guidance.
[1400] Speech recognition means: Uses technology to convert the user's voice feedback into text data.
[1401] Emotion assessment means: It has technology to assess emotions by analyzing the user's tone of voice, facial expressions, and posture.
[1402] System operation explanation
[1403] When a user puts on the AR glasses and turns them on, the camera and microphone are activated and in standby mode. When the user brings the device to be operated (e.g., a smartphone, tablet, or household appliance) into view, the camera captures the device's user interface. The captured image and audio data are sent to a cloud server via a communication means.
[1404] The cloud server analyzes the received image and audio data using an analysis means and uses a generative artificial intelligence model to analyze the user interface and audio. The emotion evaluation means also analyzes the voice tone, facial expression, and posture to evaluate the user's emotions. Based on the analysis results and the emotion evaluation, specific next steps are generated, converted into audio files and augmented reality display data, and then sent to the device.
[1405] The device provides the user with operational instructions via voice guidance and augmented reality display. For example, the voice guidance might say, "Tap the Settings app," and the location of the Settings app is highlighted on the AR glasses' display. If the user is unsure of the next step, they can ask into the microphone, "What should I do next?". The emotion assessment means analyzes the user's emotions along with this question, and the cloud server generates and provides detailed, gentle-toned instructions.
[1406] Specific examples
[1407] Example 1: Smartphone Wi-Fi settings
[1408] The user puts on the AR glasses and projects the smartphone's home screen with the camera. The device sends this image to a cloud server, which then uses generative AI to analyze the image and generate the instruction, "Tap the Settings app." The AR glasses then present this instruction to the user through voice guidance and an augmented reality display. If the user is unsure of the next step, they can ask aloud, "What should I do next?" If the emotion engine detects the user's frustration along with this question, the server generates and provides more detailed instructions in a gentler tone.
[1409] Example 2: Configuring Home Appliances
[1410] The user puts on the AR glasses and brings the control panel of a household appliance into view. The device sends an image of this panel to the server, which analyzes it. For example, an instruction such as "Press the power button" is generated and presented to the user through a guide means. If the user is unsure and gives feedback such as "This doesn't work," the emotion engine will detect anxiety and the server will generate an instruction in an encouraging tone such as "Don't worry. Next, try..."
[1411] Prompt Sentence Examples
[1412] "Please provide audio guidance and AR displays to guide users through the steps they need to set up Wi-Fi on their smartphones."
[1413] The above system aims to use a generative AI model and an emotion engine to enable elderly users to easily operate complex digital devices, and to provide flexible assistance that takes into account emotions during operation.
[1414] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1415] Step 1:
[1416] The user puts on the AR glasses.
[1417] Input: The user puts on the AR glasses and turns them on.
[1418] Operation: The camera and microphone of the AR glasses are activated and enter standby mode.
[1419] Output: The camera and microphone are ready to use.
[1420] Step 2:
[1421] Keep the device the user is controlling within sight.
[1422] Input: The user brings the device they want to control (e.g., a smartphone or household appliance) into view.
[1423] Operation: A camera means captures the device's user interface in real time.
[1424] Output: Captured image and audio data.
[1425] Step 3:
[1426] Send the data to a cloud server.
[1427] Input: Captured image and audio data.
[1428] Operation: The device compresses and encodes image and audio data and sends them to the cloud server using a communication method.
[1429] Output: The cloud server receives the captured image and audio data.
[1430] Step 4:
[1431] A cloud server analyzes the image and audio data.
[1432] Input: Image data and audio data received by the cloud server.
[1433] Operation: The analysis means (generative AI model) analyzes image data and recognizes the user interface. The emotion evaluation means analyzes voice tone, facial expressions, and posture to evaluate the user's emotions.
[1434] Output: User interface analysis results and user emotion evaluation results.
[1435] Step 5:
[1436] Generate operation guide.
[1437] Input: User interface analysis results and user emotion evaluation results.
[1438] How it works: The cloud server uses the generative AI model to create specific operating procedures in text format based on the analysis results and emotional evaluation.
[1439] Output: Generated operating instructions text.
[1440] Step 6:
[1441] The device provides operation guidance.
[1442] Input: The generated operating instructions text.
[1443] How it works: The device converts the instruction text into an audio file and augmented reality display data, and presents it to the user. Specifically, the audio guide says "Tap the Settings app," and the AR display highlights the location of the Settings app.
[1444] Output: Audio guide and augmented reality display presented to the user.
[1445] Step 7:
[1446] Process user feedback.
[1447] Input: The user interacts with the device through instructions and uses voice to ask questions or provide feedback, for example, "What should I do next?"
[1448] What it does: The device captures the user's voice feedback.
[1449] Output: Captured audio feedback data.
[1450] Step 8:
[1451] Performs voice recognition and emotion assessment.
[1452] Input: Captured audio feedback data.
[1453] Operation: The device uses a voice recognition means to convert voice data into text, and an emotion assessment means analyzes the user's voice tone and facial expressions to assess their emotions.
[1454] Output: Feedback and sentiment evaluation results converted into text.
[1455] Step 9:
[1456] Send the feedback to a cloud server.
[1457] Input: Feedback and sentiment rating results converted to text.
[1458] Operation: The device sends text data and emotion information to the cloud server using a communication method.
[1459] Output: The cloud server receives the feedback data and emotion information.
[1460] Step 10:
[1461] The cloud server generates a response.
[1462] Input: Feedback data and emotional information.
[1463] How it works: The server uses a generative AI model to generate appropriate responses based on user feedback, and flexibly adjusts the content and presentation of the response based on emotional assessment, for example, generating a calmer tone of response if the user is annoyed.
[1464] Output: The generated response text.
[1465] Step 11:
[1466] The terminal provides a response.
[1467] Input: The generated response text.
[1468] How it works: The device converts the response text into an audio file and augmented reality display data, and presents them to the user. Specifically, the audio prompt says "Select Wi-Fi settings" and the corresponding part is highlighted in the AR display.
[1469] Output: Audio guide and augmented reality display presented to the user.
[1470] Through the above processing steps, this system can provide appropriate and flexible support for operating digital devices while taking into account the user's emotions.
[1471] (Application example 2)
[1472] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1473] Elderly users face challenges when operating complex digital devices or selecting products in physical stores, such as difficulty obtaining and understanding information about operation and purchasing. Additionally, elderly users often experience anxiety and frustration while operating devices, which can be difficult to deal with appropriately. This can make it difficult for elderly users to smoothly navigate digital devices and physical stores.
[1474] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1475] In this invention, the server includes a camera, a communication unit, an analysis unit using a generative AI model, an emotion evaluation unit, a guide unit, a dialogue unit, and a voice recognition unit. This allows elderly users to easily obtain and understand product information on digital devices and in physical stores using AR glasses, and also allows them to receive flexible operation guidance that takes the user's emotions into consideration.
[1476] "AR glasses" are devices worn by elderly users that can overlay digital information onto their field of vision.
[1477] "Camera means" refers to a device for capturing an image of the user interface of a device within the user's field of view.
[1478] "Communication means" is a function for transmitting captured image data to a cloud server.
[1479] A "generative AI model" is an artificial intelligence model that analyzes received image data and generates appropriate operating procedures.
[1480] "Analysis means" is a function for analyzing transmitted images using a generative AI model.
[1481] The "guide means" is a function that provides users with operating procedures generated based on the analysis results through audio guidance and AR displays.
[1482] The "emotion evaluation means" is a function for analyzing the user's voice and facial expressions and evaluating the user's emotional state.
[1483] "Interactive means" is a function that accepts voice feedback from the user and provides additional operational guidance and emotion-sensitive guidance based on that feedback.
[1484] "Voice recognition means" is a function for converting the user's voice feedback into text data.
[1485] This invention is a system that makes it easier for elderly users to select products using complex digital devices or in physical stores. The system includes AR glasses worn by the elderly user, a camera means, a communication means, an analysis means having a generative AI model, an emotion evaluation means, a guidance means, a dialogue means, and a voice recognition means.
[1486] System Configuration
[1487] 1. AR Glasses:
[1488] This device is worn by the user and has a built-in camera, microphone, and speaker, allowing digital information to be displayed overlaid on the user's field of vision when operating a digital device or selecting products in a physical store.
[1489] 2. Camera means:
[1490] It is built into AR glasses and captures images of the device's user interface and products in physical stores within the user's field of view.
[1491] 3. Means of communication:
[1492] The captured image and audio data are sent to a cloud server using a secure, low-latency protocol.
[1493] 4. Analytical means with generative AI models:
[1494] The image data received by the server is analyzed and appropriate operating procedures and product information are generated. The generative AI model learns from a large data set and provides highly accurate analysis results.
[1495] 5. Emotional assessment measures:
[1496] It analyzes the user's tone of voice, facial expressions, and posture to assess the user's emotional state, allowing it to process user feedback appropriately and provide emotionally sensitive assistance.
[1497] 6. Guide means:
[1498] Based on the analysis results, the system generates specific next steps and product information, and provides these instructions to the user in the form of audio guidance or AR displays.
[1499] 7. Means of interaction:
[1500] If the user is unsure of a procedure or needs additional information, the system accepts voice feedback and provides additional operating instructions or emotionally sensitive guidance based on the feedback.
[1501] 8. Voice Recognition Methods:
[1502] The user's voice feedback is converted into text data, which is further analyzed by generative AI models and sentiment assessment methods.
[1503] Example
[1504] Example 1: Smartphone settings
[1505] The user puts on the AR glasses and projects the smartphone's home screen onto the camera. The cloud server analyzes this image using a generative AI model and provides the instruction, "Tap the Settings app." The AR glasses present this instruction to the user through voice guidance and AR displays. If the user is unsure of the next step, they can ask aloud, "What should I do next?" If the emotion engine detects the user's frustration along with this question, the server generates and provides more detailed instructions in a gentler tone.
[1506] Example 2: Product selection in a store
[1507] A user enters a supermarket and wants to look at a particular shelf. The AR glasses send an image of the shelf to a cloud server. The cloud server analyzes the image using a generative AI model and provides a voice prompt saying, "This product is rich in probiotics." If the user's emotions are unstable, the emotion assessment mechanism detects this and provides instructions in an encouraging tone, such as, "Don't worry, check out this product next."
[1508] Prompt example
[1509] "What should I do next?"
[1510] "Is this the right button to press?"
[1511] As described above, this system allows elderly users to easily operate digital devices and select products in physical stores. Furthermore, the emotional support allows them to use the system with peace of mind.
[1512] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1513] Step 1:
[1514] A user wears AR glasses and places a digital device or product in their field of view. The camera built into the AR glasses captures an image of the user interface or product. This image data becomes the input for the system. The camera means performs real-time image capture.
[1515] Step 2:
[1516] The device sends image data captured by the camera to the cloud server via a communication means. At this time, the image data is compressed, encoded, and sent with low latency. Data transmission is performed safely and quickly using a communication protocol.
[1517] Step 3:
[1518] The server analyzes the received image data using a generative AI model. This analysis method identifies the user interface and product attributes in the image and generates operation procedures and product information as the analysis results. The input is image data, and the output is the analysis results.
[1519] Step 4:
[1520] Based on the analysis results, the server generates the next operation procedure and product information, which the guide means provides to the terminal in the form of audio guide and AR display. The guide means converts the generated text into an audio file and AR display data.
[1521] Step 5:
[1522] The user operates the device and checks the product according to the presented operating instructions and product information. If the user is unsure of the next step, voice feedback is provided through the device's microphone. A prompt such as "What should I do next?" is input.
[1523] Step 6:
[1524] The device captures the user's voice feedback and converts it into text data using a voice recognition means. At the same time, the emotion assessment means analyzes the user's voice tone and facial expressions to evaluate the user's emotional state. The input is voice feedback, and the output is text data and emotion information.
[1525] Step 7:
[1526] The text data and emotional information converted by the speech recognition means are sent to a cloud server. The server then inputs this data back into the generative AI model to generate an appropriate response. The content and presentation of the response are flexibly adjusted based on the analysis results of the emotional evaluation means.
[1527] Step 8:
[1528] The server converts the generated response into an audio file and AR display data and sends them to the terminal. Based on this, the guide means provides the user with a response again in the form of audio guidance and AR display. For example, an instruction such as "Please select Wi-Fi settings" is generated.
[1529] This series of steps allows elderly users to smoothly operate digital devices and select products in physical stores.
[1530] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1531] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1532] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1533] [Fourth embodiment]
[1534] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1535] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1536] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1537] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1538] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1539] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1540] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1541] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1542] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1543] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1544] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1545] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1546] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1547] The present invention relates to a support system for elderly people to utilize generative AI and IT tools. This system uses an AR glasses camera to image-recognize the user interface of the device operated by the user, and analyzes the image using a generative AI model to provide operational support using audio guidance and AR displays. The following describes in detail the embodiments of the present invention.
[1548] System Configuration
[1549] The system mainly includes the following components:
[1550] AR glasses: Worn by elderly users, these glasses have built-in cameras, microphones, and speakers.
[1551] Camera means: Built into the AR glasses, it captures the user interface of the device within the user's field of view.
[1552] Communication method: Captured image data is sent to a cloud server.
[1553] Cloud server: Receives image data and has a generative AI model that performs analysis.
[1554] Analysis method: Analyze the user interface in the image using a generative AI model.
[1555] Guidance means: Operating procedures are generated based on the analysis results and presented to the user through audio guidance and AR displays.
[1556] Interaction methods: Receives user voice feedback and generates additional operation instructions based on the feedback.
[1557] Speech recognition means: converts the user's voice feedback into text data.
[1558] Program processing flow
[1559] 1. The user puts on the AR glasses
[1560] The user puts on the AR glasses and turns on the device, which activates the camera, microphone, and other components.
[1561] 2. User Operation
[1562] The user places the device (e.g., smartphone, tablet, or household appliance) they want to operate within their field of view, and the camera captures the user interface of this device in real time.
[1563] 3. Data Transmission
[1564] The terminal transmits the captured image data to a cloud server using a communication means.
[1565] 4. Cloud server analysis
[1566] The server analyzes the received image data using an analysis means and generates device operation instructions using a generative AI model, such as "Tap the Settings app."
[1567] 5. Provision of guides
[1568] The server converts the generated operating procedures into audio files and AR display data and sends them to the terminal.
[1569] Based on this, the device will use voice guidance and AR displays to guide the user on how to operate the device.
[1570] 6. User Feedback
[1571] If the user is confused about an operation, the device will provide voice feedback through the microphone, for example, asking, "What should I do next?"
[1572] 7. Voice Recognition and Further Guidance
[1573] The terminal converts the voice feedback into text using a voice recognition means and transmits it to a cloud server.
[1574] The server analyzes this feedback, generates additional operating instructions, and again uses the guide means to provide operating assistance to the user.
[1575] Specific examples
[1576] Example 1: Smartphone Wi-Fi settings
[1577] The user puts on the AR glasses and projects the smartphone's home screen with the camera. The device sends this image to a cloud server, which then uses generative AI to analyze the image and generate the instruction, "Tap the Settings app." The AR glasses then present this instruction to the user through audio guidance and AR displays. If the user is unsure of the next step, they can ask, "What should I do next?" and the server will continue to generate and provide additional instructions.
[1578] Example 2: Configuring Home Appliances
[1579] The user puts on the AR glasses and places the control panel of a household appliance in their field of view. The device sends an image of this panel to the server, which analyzes it. For example, an instruction such as "Press the power button" is generated and presented to the user through the guidance means. If the user is still unsure about how to operate the device, assistance continues through similar voice feedback and analysis.
[1580] In accordance with the above aspects, the present invention realizes a system that enables elderly people to easily understand and use the operation of complex digital devices.
[1581] The processing flow will be explained below.
[1582] Step 1:
[1583] The user puts on the AR glasses and turns on the device, which activates the camera and microphone and puts them into standby mode.
[1584] Step 2:
[1585] The user brings the device to be operated (e.g., a smartphone or a household appliance) into view, and the camera means captures the user interface of the device in real time.
[1586] Step 3:
[1587] The device compresses and encodes the captured images and sends them to a cloud server with low latency.
[1588] Step 4:
[1589] The server decodes the image data sent and inputs it into a generative artificial intelligence model (generative AI) to begin analysis.
[1590] Step 5:
[1591] A generative AI built into the server analyzes UI elements in the image and identifies the device's current state (e.g., "home screen" or "settings screen").
[1592] Step 6:
[1593] Based on the analysis results, the server generates specific instructions in text format for the next steps, such as "Tap the settings icon."
[1594] Step 7:
[1595] The operating procedures generated by the server are converted into audio files and AR display data and sent to the terminal.
[1596] Step 8:
[1597] Based on the operation guide received by the device from the server, the device presents the user with operation procedures using voice guidance and AR display. For example, the voice guidance may say "Tap the Settings app" and the AR display may show the location of the Settings app.
[1598] Step 9:
[1599] The user operates the device according to the operating instructions. If the user is unsure of an operation, they can ask questions or provide feedback by speaking into the microphone. For example, "What should I do next?"
[1600] Step 10:
[1601] The device captures the user's voice and converts it into text data using voice recognition technology.
[1602] Step 11:
[1603] The terminal transmits the converted text data to the server.
[1604] Step 12:
[1605] The server then re-feeds the user's questions and feedback into the generative AI model to generate an appropriate response, such as "Select your Wi-Fi settings."
[1606] Step 13:
[1607] The server converts the response into an audio file or AR display data and sends it to the device.
[1608] Step 14:
[1609] The device will again provide the response to the user through audio guidance and AR display.
[1610] Step 15:
[1611] The user resumes operation according to further instructions from the server, repeating steps 9 through 14 as necessary.
[1612] Example 1
[1613] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1614] Currently, the difficulty that older people experience when operating complex electronic devices and software is a major problem. Understanding new operation methods and interfaces is a challenge, especially when introducing new technologies and tools, creating technical barriers. To address this issue, assistance systems that can be operated intuitively by older people are needed, and the importance of real-time operation guidance and feedback systems is particularly increasing.
[1615] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1616] In this invention, the server includes a camera, a communication device, and an analysis device that uses a generative AI model. This enables elderly users to operate electronic devices quickly and accurately through real-time, intuitive voice guidance and AR displays. Specifically, the server captures images of the user's operating environment with a camera, sends the images to the server, and analyzes them with a generative AI model to provide appropriate operation guidance. Furthermore, the server converts the user's voice feedback into text for further operational assistance, making it easier to operate complex digital devices.
[1617] "Elderly users" refers to older individuals who are the primary users of the system and who have difficulty using electronic devices and software, especially those that are complex to operate.
[1618] "AR glasses" are headset-type devices that users can wear and use, equipped with augmented reality (AR) technology to display information as an overlay on their field of vision.
[1619] "Photographing means" refers to the camera built into the AR glasses, a device that has the function of capturing an image of the user interface of an electronic device within the user's field of view.
[1620] "Communication means" refers to data transmission devices such as Wi-Fi modules and 4G / 5G communication modules for transmitting captured image data to a remote server.
[1621] A "generative artificial intelligence model" is an AI model used for image recognition and natural language processing, and refers to a program that generates specific operating instructions and guides based on deep learning technology, for example.
[1622] "Analysis means" refers to software and hardware for analyzing the transmitted image data and recognizing the state of the electronic device, the position of the user interface, and the like.
[1623] "Operation guidance means" refers to a mechanism for providing users with operation procedures generated based on the analysis results through audio guidance and AR displays.
[1624] "Interactive means" refers to a device or program that has the function of receiving voice feedback from a user and generating and providing additional operation guidance based on that feedback.
[1625] "Speech recognition means" means the technology or software used to convert a user's voice feedback into text data.
[1626] "Presentation means" refers to a screen display or audio output device that visualizes and presents to the user the operation instructions generated based on the analyzed data.
[1627] This invention relates to a support system that enables elderly people to easily operate complex electronic devices. This system includes a photographing means, a communication means, an analysis means using a generative artificial intelligence model, an operation guidance means, a dialogue means, a voice recognition means, and a presentation means. The following describes in detail an embodiment of the invention.
[1628] Hardware and Software Configuration
[1629] Hardware used
[1630] 1. AR Glasses:
[1631] Built-in camera: A camera for capturing the user interface of an electronic device in the user's field of view.
[1632] Built-in microphone: A microphone for collecting user voice feedback.
[1633] Built-in speaker: A speaker for playing audio guides.
[1634] Communication module: A module that supports Wi-Fi and 4G / 5G communications.
[1635] 2. Cloud Server:
[1636] high performance computer
[1637] Generative AI models (e.g., GPT-4)
[1638] Image Analysis Software
[1639] Software used
[1640] 1. Real-time image capture software: Software that captures images from the AR glasses camera and sends them to a cloud server.
[1641] 2. Speech recognition software: Software for converting user voice feedback into text data.
[1642] 3. Generative AI model: Software that analyzes the transmitted image data and generates appropriate operating instructions.
[1643] 4. Speech synthesis software: Software for converting the generated operating instructions into audio guidance.
[1644] 5. AR display software: Software for overlaying operation guides onto the user's field of view.
[1645] System Operation
[1646] Example 1: Smartphone Wi-Fi settings
[1647] Consider a scenario where a user wears AR glasses and operates the home screen of a smartphone. When the user looks at the smartphone, the camera in the AR glasses captures the home screen, and the image data is sent to a cloud server. The cloud server analyzes this image using a generative AI model and generates an operation instruction such as "Tap the Settings app." This instruction is converted into an audio file and AR display data, which are then sent to the device. The user follows these instructions to perform the operation.
[1648] As a concrete example, the following prompt sentences could be input to a generative AI model:
[1649] "I want to set up Wi-Fi on my smartphone. Can you please tell me the detailed steps to tap the Settings app?"
[1650] Example 2: Home electronics settings
[1651] Consider a scenario where a user wears AR glasses and operates a home electronic device. For example, let's say they want to operate an air conditioner remote control. When the user picks up the remote control, the camera in the AR glasses captures the control panel, and the image is sent to a cloud server. After analysis by the generative AI model, an instruction such as "Press the power button" is generated. This instruction is provided to the user through audio guidance and AR displays.
[1652] As a concrete example, the following prompt sentences could be input to a generative AI model:
[1653] "I want to set up a home appliance. First, please tell me the steps to press the power button."
[1654] As described above, this system helps seniors easily understand and operate complex electronic devices. By providing real-time operation guidance and additional guidance based on user feedback, it is possible to reduce technological barriers and promote the use of digital devices.
[1655] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1656] Step 1:
[1657] The user puts on the AR glasses and turns on the device.
[1658] Specific operation: The user presses and holds the power button on the AR glasses to start the system, which automatically turns on the camera, microphone, speaker, etc.
[1659] Input: User action (pressing the power button)
[1660] Output: Notification of AR glasses startup status and readiness
[1661] Step 2:
[1662] Keep the device the user wants to control within sight.
[1663] Specific Action: A user brings an electronic device, such as a smartphone or consumer electronic device, into their field of view. The camera in the AR glasses captures this user interface.
[1664] Input: Placing the device in view
[1665] Output: Captured image data of the user interface
[1666] Step 3:
[1667] The device transmits the captured image data to a cloud server.
[1668] Specific operation: The device's communication means sends the captured image data to the cloud server via Wi-Fi or 4G / 5G. The user is notified of the progress during the transmission.
[1669] Input: Captured image data
[1670] Output: Notification of completion of data transmission to the cloud server
[1671] Step 4:
[1672] The server analyzes the received image data.
[1673] Specific operation: The server begins analyzing the received image data using analytical means. A generative AI model (e.g., GPT-4) is used for the analysis to recognize the device state and interface.
[1674] Input: Image data sent
[1675] Output: Parsed data (current state of the device and possible actions)
[1676] Step 5:
[1677] The server generates operation instructions using the generative AI model.
[1678] Specific actions: The generative AI model generates appropriate instructions based on the analysis results. For example, it creates specific instructions such as "Tap the Settings app."
[1679] Input: Parsed data
[1680] Output: Generated operating instructions
[1681] Step 6:
[1682] The operating procedures generated by the server are converted into audio files and AR display data.
[1683] Specific operation: The server converts the generated operation instructions into a voice file and AR display data using voice synthesis software and AR display software.
[1684] Input: Generated operation instructions
[1685] Output: Audio file and AR display data
[1686] Step 7:
[1687] The server sends the generated audio files and AR display data to the terminal.
[1688] Specific operation: The server sends the audio file and AR display data to the device using a communication method. The user is notified of the transmission status.
[1689] Input: Audio files and AR display data
[1690] Output: Notification of completion of transmission to the terminal
[1691] Step 8:
[1692] The device will guide the user on how to operate the device through voice guidance and AR displays.
[1693] Specific operation: The device plays the received audio file through the speaker and displays an overlay of operation instructions on the AR glasses display, such as "Tap the Settings app."
[1694] Input: Audio files and AR display data
[1695] Output: User guide
[1696] Step 9:
[1697] If the user has difficulty operating the device, voice feedback is provided.
[1698] Specific Action: If a user is unsure about an action, they can speak into the microphone and ask a question like, "What should I do next?" The microphone will capture their voice.
[1699] Input: User's voice feedback
[1700] Output: Captured audio data
[1701] Step 10:
[1702] The terminal converts the voice feedback into text using a voice recognition means and transmits it to a cloud server.
[1703] Specific operation: The device converts the voice data into text data using voice recognition software and sends the text data to the cloud server.
[1704] Input: Captured audio data
[1705] Output: Text data and notification of completion of transmission to the cloud server
[1706] Step 11:
[1707] The server analyzes the text data and generates additional operating instructions.
[1708] Specific operation: The server analyzes the received text data and generates additional operation instructions based on the user's question, such as "Next, select Wi-Fi."
[1709] Input: Text data
[1710] Output: Additional operating instructions
[1711] Step 12:
[1712] The server transmits the generated additional operation instructions to the terminal, and the terminal provides the user with additional guidance.
[1713] Specific operation: The server converts additional operation instructions into an audio file and AR display data and sends them to the terminal, which then provides guidance to the user based on this.
[1714] Input: Additional operating instructions
[1715] Output: Operation guide with additional audio files and AR display data
[1716] The above is the flow of processing in the program for this system and the specific operations of each.
[1717] (Application example 1)
[1718] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1719] When elderly people make electronic payments using smartphones and other digital devices, the operation is complicated and difficult for them to use. In addition, there is a high possibility that elderly people will make operational errors or face security risks, so an efficient method to support them is needed.
[1720] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1721] In this invention, the server includes an analysis means, an operation guidance means, and a dialogue means, which allows elderly users to be guided through the transaction procedure through the AR glasses when making electronic payments using their smartphones, and provides specific operation instructions according to the progress of the transaction.
[1722] "Elderly users" refer to users who are older and unfamiliar with operating digital devices.
[1723] "Wearable AR glasses" are glasses-type devices that users wear on their heads to display augmented reality.
[1724] A "user interface" refers to the part of a device that a user directly operates or inputs, such as a screen or operation panel.
[1725] "Imaging means" refers to the function of capturing images using a camera or sensor.
[1726] "Data communication means" refers to the function of transmitting captured data to other devices or servers.
[1727] A "cloud computing server" refers to a remote server used to store and analyze data over the Internet.
[1728] A "generative artificial intelligence model" is an artificial intelligence that analyzes input and data from users and generates appropriate responses and instructions.
[1729] "Analysis means" refers to a function that analyzes the captured image data and extracts information about the user's operations.
[1730] "Operation guide means" refers to a function that presents specific operation methods to the user based on the analysis results.
[1731] "Augmented reality display" refers to a technology that adds virtual information to the real field of view and displays it.
[1732] "Interaction" refers to the ability to interact with the user and provide additional guidance based on their feedback.
[1733] "Electronic payment support means" refers to a function that supports operations during electronic payment and guides transaction procedures.
[1734] "Transaction guidance means" refers to a function that provides specific operational instructions, such as scanning a QR code or tapping a payment button, depending on the progress of the transaction.
[1735] The system for realizing this invention is designed to help elderly users smoothly make electronic payments. The system is realized by the following components and operations.
[1736] System Configuration
[1737] The system mainly includes the following components:
[1738] 1. AR glasses: A device worn by elderly users that has a built-in camera, microphone, and speaker.
[1739] 2. Imaging means: The camera built into the AR glasses captures the device's user interface in real time as it comes into the user's field of view.
[1740] 3. Data communication means: Transmit the captured image data to a cloud computing server.
[1741] 4. Cloud computing server: Analyzes the received image data and generates operation guides using a generative AI model.
[1742] 5. Analysis method: Analyze the user interface in the image on the cloud server and identify the next operation to be performed.
[1743] 6. Operation guide means: Based on the analysis results, operation procedures are generated as audio guides and AR displays and provided to the user.
[1744] 7. Interaction: Receives voice feedback from the user and provides additional guidance based on that.
[1745] 8. Speech recognition means: converts the user's voice feedback into text data.
[1746] 9. Electronic payment support measures: Provide specific operating procedures for electronic payments.
[1747] 10. Transaction Guidance: Guidance on transaction operations such as scanning a QR code or tapping a payment button.
[1748] How it works
[1749] 1. Function of the imaging means: When a user wears the AR glasses and tries to operate a smartphone, the camera captures the smartphone screen and obtains image data in real time.
[1750] 2. Function of data communication means: The acquired image data is transmitted to the cloud computing server via data communication means (e.g., Wi-Fi or mobile network).
[1751] 3. Function of the analysis tool: A generative AI model on a cloud server analyzes image data, for example, identifying QR codes or specific buttons.
[1752] 4. Operation guidance function: Based on the analysis results, the cloud server generates specific operation instructions, such as "Scan the QR code" or "Tap the payment button."
[1753] 5. Audio guide and AR display: The generated operating procedure is provided to the user as an audio guide and an augmented reality display. The next operation the user should perform is presented visually and audibly.
[1754] 6. Interaction function: If the user is unsure of what to do, they can provide voice feedback (e.g., "Which button should I press next?"). This feedback is sent to the server via the microphone.
[1755] 7. Voice recognition and additional guidance: Voice feedback is converted into text data by a voice recognition tool and analyzed again by the cloud server. Additional specific operating instructions are then generated and provided again as voice guidance or AR displays.
[1756] Specific examples
[1757] Example 1: Payment processing for an electronic payment app
[1758] The user puts on the AR glasses and opens the electronic payment app. The smartphone screen is captured and analyzed by the cloud server. The server generates an instruction to "scan the QR code" and displays it as an audio guide and AR display through the AR glasses. When the user follows the instructions, the next instruction is displayed: "tap the payment button."
[1759] Prompt: A user is looking at an electronic payment app. Guide them through the next steps in the payment process.
[1760] Instructions for the generative AI model: "Scan the QR code" or "Tap the payment button"
[1761] Example 2: Feedback and additional guidance
[1762] The user is confused and asks, "Which button should I press next?" This voice input is converted into text data and sent to a cloud server. The server analyzes the feedback and generates additional instructions as the next step: "Tap the back button and reselect your payment method."
[1763] This invention enables elderly people to easily understand how to operate complex digital devices and use them safely.
[1764] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1765] Step 1:
[1766] A user puts on the AR glasses and attempts to operate a smartphone. The user launches an electronic payment app, and the camera in the AR glasses captures this screen in real time.
[1767] Input: The smartphone screen operated by the user
[1768] Output: Real-time captured image data
[1769] Specific operation: The camera in the AR glasses captures an image of the user interface and acquires it as image data.
[1770] Step 2:
[1771] The terminal transmits the acquired image data to a cloud computing server via a communication means, with the data transfer occurring via an internet connection.
[1772] Input: Image data captured in real time
[1773] Output: Image data sent to a cloud computing server
[1774] Specific operation: The communication module in the AR glasses transmits image data using Wi-Fi or a mobile network.
[1775] Step 3:
[1776] The server analyzes the received image data using an analytical method and uses a generative AI model to determine the next step to take, such as whether the user should scan a QR code or tap a payment button.
[1777] Input: Image data sent to a cloud computing server
[1778] Output: Specific operation procedures as analysis results
[1779] Specific operation: The generative artificial intelligence model analyzes the image data and generates appropriate operating instructions based on the prompt text.
[1780] Step 4:
[1781] The server generates the operating procedures as audio guidance and AR display and sends them to the terminal.
[1782] Input: Specific operation procedures as analysis results
[1783] Output: Audio guide and AR display data
[1784] Specific operation: The operation guide generation module in the server converts the instructions into audio files and AR display data.
[1785] Step 5:
[1786] The device uses the received voice guidance and AR display to guide the user on how to operate the device.
[1787] Input: Audio guide and AR display data
[1788] Output: Audio guidance and AR display to the user
[1789] What it does: Uses the speakers and display devices of the AR glasses to provide audio and visual guidance.
[1790] Step 6:
[1791] If the user is confused about an operation, they can receive voice feedback, such as asking, "Which button should I press next?" This voice feedback is collected through the microphone.
[1792] Input: User's voice feedback
[1793] Output: Collected audio data
[1794] Specific operation: The microphone inside the AR glasses captures sound and saves it as audio data.
[1795] Step 7:
[1796] The terminal converts the voice feedback into text using a voice recognition means and transmits it to a cloud server.
[1797] Input: Collected audio data
[1798] Output: Feedback data converted to text
[1799] Specific operation: Uses voice recognition technology to convert voice data into text data.
[1800] Step 8:
[1801] The server analyzes the text feedback and generates additional operating instructions, which are then provided to the user again as audio guidance and AR displays.
[1802] Input: Feedback data converted to text
[1803] Output: Audio guide and AR display data as additional operating instructions
[1804] Specific behavior: The generative AI model analyzes the feedback, generates new operating instructions, and presents them to the user again.
[1805] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1806] The present invention combines an emotion engine with a support system for elderly users to utilize generative AI. This system uses an AR glasses camera to image-recognize the user interface of the device operated by the user, and analyzes the image using a generative AI model to provide operational support using audio guidance and AR displays, as well as flexible guidance by recognizing the user's emotions. The following describes in detail the embodiments of the present invention.
[1807] System Configuration
[1808] The system mainly includes the following components:
[1809] AR glasses: Worn by elderly users, these glasses have built-in cameras, microphones, and speakers.
[1810] Camera means: Built into the AR glasses, it captures the user interface of the device within the user's field of view.
[1811] Communication means: Captured image data and audio data are sent to a cloud server.
[1812] Cloud server: Receives image data and audio data and has a generative AI model that analyzes them.
[1813] Analysis method: A generative AI model is used to analyze the user interface and audio data in the image.
[1814] Guidance means: Based on the analysis results, operating procedures are generated and presented to the user through audio guidance and AR displays.
[1815] Interaction methods: Accepts user voice feedback and provides additional guidance based on the feedback.
[1816] Speech recognition means: converts the user's voice feedback into text data.
[1817] Emotion Engine: Evaluates emotions by analyzing the user's tone of voice, facial expressions, and posture.
[1818] Program processing flow
[1819] 1. The user puts on the AR glasses
[1820] The user puts on the AR glasses and turns them on, which activates the camera and microphone and puts them into standby mode.
[1821] 2. User Actions
[1822] The user brings the device to be operated (e.g., a smartphone or a household appliance) into view, and the camera means captures the user interface of this device in real time.
[1823] 3. Data Transmission
[1824] The image and audio data captured by the device is compressed and encoded and sent to a cloud server with low latency.
[1825] 4. Cloud server analysis
[1826] The image and audio data received by the server is analyzed by an analysis means, and the user interface and audio are analyzed using a generative artificial intelligence model (generative AI).
[1827] The emotion engine analyzes voice tone, facial expressions, and posture to assess the user's emotions.
[1828] 5. Providing operation guides
[1829] Based on the analysis results and emotion evaluation, the server converts the specific next steps into text format, converts it into an audio file or AR display data, and sends it to the device.
[1830] Based on this, the device will present the user with operation instructions using voice guidance and AR display. For example, the voice guidance will say "Tap the Settings app" and the AR display will show the location of the Settings app.
[1831] 6. User Feedback
[1832] The user operates the device by following the instructions. If they are unsure of how to operate the device, they can ask questions or provide feedback by speaking into the microphone. For example, "What should I do next?"
[1833] 7. Speech Recognition and Emotion Assessment
[1834] The device captures the user's voice feedback and converts it into text data using a speech recognition tool, while the emotion engine evaluates the user's emotions.
[1835] 8. Submitting Feedback
[1836] The device sends the converted text data and emotional information to a cloud server.
[1837] 9. Response generation by cloud server
[1838] The server then re-feeds the user's questions and feedback into the generative AI model to generate an appropriate response, such as "Select your Wi-Fi settings."
[1839] Based on information from the emotion engine, the content and presentation of the response can be flexibly adjusted, for example, if the user is annoyed, a calmer tone of response can be generated.
[1840] 10. Providing a Response
[1841] The server converts the response into an audio file or AR display data and sends it to the device.
[1842] The device will again provide the response to the user through audio guidance and AR display.
[1843] Specific examples
[1844] Example 1: Smartphone Wi-Fi settings
[1845] The user puts on the AR glasses and projects the smartphone's home screen with the camera. The device sends this image to a cloud server, which then uses generative AI to analyze the image and generate the instruction, "Tap the Settings app." The AR glasses then present this instruction to the user through voice guidance and AR display. If the user is unsure of the next step, they can ask aloud, "What should I do next?" If the emotion engine detects the user's frustration along with this question, the server generates and provides more detailed instructions in a gentler tone.
[1846] Example 2: Configuring Home Appliances
[1847] The user puts on the AR glasses and brings the control panel of a household appliance into view. The device sends an image of this panel to the server, which analyzes it. For example, an instruction such as "Press the power button" is generated and presented to the user through a guide means. If the user is unsure and gives feedback such as "This doesn't work," the emotion engine will detect anxiety and the server will generate an instruction in an encouraging tone such as "Don't worry. Next, try..."
[1848] In accordance with the above aspects, the present invention realizes a system that enables elderly people to easily understand and use the operation of complex digital devices, and provides flexible support that takes emotions into consideration.
[1849] The processing flow will be explained below.
[1850] Step 1:
[1851] The user puts on the AR glasses, which activates the camera and microphone of the AR glasses and puts them into standby mode.
[1852] Step 2:
[1853] The user brings the device to be operated (e.g., a smartphone or a home appliance) into view. The device uses a camera to capture the device's user interface in real time.
[1854] Step 3:
[1855] The device compresses and encodes the captured image data and user voice data and sends them to the cloud server.
[1856] Step 4:
[1857] The server decodes the received image data, and the analysis means analyzes the user interface elements in the image using a generative artificial intelligence model (generative AI).
[1858] Step 5:
[1859] The server analyzes the received voice data and converts it into text using a voice recognition means.
[1860] Step 6:
[1861] The server uses the generation AI to generate the next steps based on the analysis results, creating specific instructions such as "tap the Settings app."
[1862] Step 7:
[1863] The server converts the generated operating procedures into audio files and AR display data and sends them to the terminal.
[1864] Step 8:
[1865] Based on the operation guide received by the device from the server, the device presents the user with operation procedures using voice guidance and AR display. For example, the voice guidance may say "Tap the Settings app" and the AR display may show the location of the Settings app.
[1866] Step 9:
[1867] The user operates the device according to the operating instructions. If the user is unsure of how to operate the device, they can ask questions into the microphone, such as "What should I do next?"
[1868] Step 10:
[1869] The device captures the user's voice feedback and converts it into text data using voice recognition technology, while the emotion engine analyzes the user's voice tone, facial expressions, and posture to make an emotional assessment.
[1870] Step 11:
[1871] The device sends the converted text data and emotional information to a cloud server.
[1872] Step 12:
[1873] The server re-inputs the user's questions and feedback into the AI model to generate an appropriate response. Based on the evaluation of the emotion engine, the model generates instructions with a tone and content that corresponds to the user's emotions. For example, if the user is showing signs of irritation, the model generates instructions in a calm tone such as "Don't worry. Next, please select your Wi-Fi settings."
[1874] Step 13:
[1875] The server converts the response into an audio file or AR display data and sends it to the device.
[1876] Step 14:
[1877] The device will again provide the response to the user through audio guidance and AR display.
[1878] Step 15:
[1879] The user resumes operation following any further instructions from the server, repeating steps 9 through 14 as necessary.
[1880] Example 2
[1881] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1882] When elderly people operate complex digital devices, they often find it difficult to understand the operating procedures, and emotions such as anxiety and frustration during operation can hinder their operation. To solve this problem, it is necessary to provide appropriate operating instructions that take into account the user's emotions.
[1883] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1884] In this invention, the server includes an interaction means for receiving voice feedback from the user and providing additional operation guidance based on the feedback, a voice recognition means for converting the user's voice feedback into text using voice recognition technology, and an emotion evaluation means for analyzing the voice tone, facial expression, and posture to evaluate the user's emotions, thereby making it possible to provide flexible and appropriate operation guidance that takes the user's emotions into consideration.
[1885] "Elderly users" refer to users who are likely to find it difficult to operate digital devices due to their age.
[1886] "AR glasses" refers to a glasses-type device that uses augmented reality technology to overlay information on the real world.
[1887] "Capture means" refers to a camera or other capture device used to capture images of objects or interfaces within the user's field of view.
[1888] "Communication means" refers to an internet connection or wireless communication means for sending and receiving image and audio data between the terminal and the cloud server.
[1889] A "generative artificial intelligence model" refers to a machine learning model designed to perform a specific task (e.g., image recognition or speech analysis).
[1890] The "analysis means" refers to a digital processor that analyzes the acquired image and audio data, recognizes the user interface, and generates operation procedures.
[1891] "Guidance means" refers to an audio output device or an augmented reality display device that notifies the user of the operating procedure generated based on the analysis results.
[1892] "Interaction means" refers to the part of the system that accepts voice or other operational input from the user and generates additional operational instructions based on that input.
[1893] "Speech recognition means" refers to technology or equipment for converting voice data obtained from a user into text data.
[1894] "Emotion assessment means" refers to technology or devices for analyzing a user's tone of voice, facial expression, posture, etc. to assess the user's emotional state.
[1895] This system aims to assist elderly users in operating complex digital devices, and uses AR glasses to capture the user interface of the target device within the user's field of view. Specifically, a camera means built into the AR glasses captures images of the user device in real time and transmits them to a cloud server via a communication means.
[1896] Hardware and software used
[1897] AR glasses: Devices that display information using augmented reality technology and have built-in cameras, microphones, and speakers.
[1898] Camera means: Capable of capturing the user interface of a device within the user's field of view.
[1899] Communication means: Includes internet connection and wireless communication means for transmitting captured image and audio data to a cloud server.
[1900] Cloud server: Receives image and audio data and analyzes them using generative artificial intelligence models.
[1901] Analysis means: A generative artificial intelligence model in the cloud server is used to analyze the user interface and audio data in the image.
[1902] Guidance means: Based on the analysis results, operating procedures are generated and presented to the user through audio guidance and augmented reality displays.
[1903] Interaction methods: Accepts user voice feedback and provides additional guidance.
[1904] Speech recognition means: Uses technology to convert the user's voice feedback into text data.
[1905] Emotion assessment means: It has technology to assess emotions by analyzing the user's tone of voice, facial expressions, and posture.
[1906] System operation explanation
[1907] When a user puts on the AR glasses and turns them on, the camera and microphone are activated and in standby mode. When the user brings the device to be operated (e.g., a smartphone, tablet, or household appliance) into view, the camera captures the device's user interface. The captured image and audio data are sent to a cloud server via a communication means.
[1908] The cloud server analyzes the received image and audio data using an analysis means and uses a generative artificial intelligence model to analyze the user interface and audio. The emotion evaluation means also analyzes the voice tone, facial expression, and posture to evaluate the user's emotions. Based on the analysis results and the emotion evaluation, specific next steps are generated, converted into audio files and augmented reality display data, and then sent to the device.
[1909] The device provides the user with operational instructions via voice guidance and augmented reality display. For example, the voice guidance might say, "Tap the Settings app," and the location of the Settings app is highlighted on the AR glasses' display. If the user is unsure of the next step, they can ask into the microphone, "What should I do next?". The emotion assessment means analyzes the user's emotions along with this question, and the cloud server generates and provides detailed, gentle-toned instructions.
[1910] Specific examples
[1911] Example 1: Smartphone Wi-Fi settings
[1912] The user puts on the AR glasses and projects the smartphone's home screen with the camera. The device sends this image to a cloud server, which then uses generative AI to analyze the image and generate the instruction, "Tap the Settings app." The AR glasses then present this instruction to the user through voice guidance and an augmented reality display. If the user is unsure of the next step, they can ask aloud, "What should I do next?" If the emotion engine detects the user's frustration along with this question, the server generates and provides more detailed instructions in a gentler tone.
[1913] Example 2: Configuring Home Appliances
[1914] The user puts on the AR glasses and brings the control panel of a household appliance into view. The device sends an image of this panel to the server, which analyzes it. For example, an instruction such as "Press the power button" is generated and presented to the user through a guide means. If the user is unsure and gives feedback such as "This doesn't work," the emotion engine will detect anxiety and the server will generate an instruction in an encouraging tone such as "Don't worry. Next, try..."
[1915] Prompt Sentence Examples
[1916] "Please provide audio guidance and AR displays to guide users through the steps they need to set up Wi-Fi on their smartphones."
[1917] The above system aims to use a generative AI model and an emotion engine to enable elderly users to easily operate complex digital devices, and to provide flexible assistance that takes into account emotions during operation.
[1918] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1919] Step 1:
[1920] The user puts on the AR glasses.
[1921] Input: The user puts on the AR glasses and turns them on.
[1922] Operation: The camera and microphone of the AR glasses are activated and enter standby mode.
[1923] Output: The camera and microphone are ready to use.
[1924] Step 2:
[1925] Keep the device the user is controlling within sight.
[1926] Input: The user brings the device they want to control (e.g., a smartphone or household appliance) into view.
[1927] Operation: A camera means captures the device's user interface in real time.
[1928] Output: Captured image and audio data.
[1929] Step 3:
[1930] Send the data to a cloud server.
[1931] Input: Captured image and audio data.
[1932] Operation: The device compresses and encodes image and audio data and sends them to the cloud server using a communication method.
[1933] Output: The cloud server receives the captured image and audio data.
[1934] Step 4:
[1935] A cloud server analyzes the image and audio data.
[1936] Input: Image data and audio data received by the cloud server.
[1937] Operation: The analysis means (generative AI model) analyzes image data and recognizes the user interface. The emotion evaluation means analyzes voice tone, facial expressions, and posture to evaluate the user's emotions.
[1938] Output: User interface analysis results and user emotion evaluation results.
[1939] Step 5:
[1940] Generate operation guide.
[1941] Input: User interface analysis results and user emotion evaluation results.
[1942] How it works: The cloud server uses the generative AI model to create specific operating procedures in text format based on the analysis results and emotional evaluation.
[1943] Output: Generated operating instructions text.
[1944] Step 6:
[1945] The device provides operation guidance.
[1946] Input: The generated operating instructions text.
[1947] How it works: The device converts the instruction text into an audio file and augmented reality display data, and presents it to the user. Specifically, the audio guide says "Tap the Settings app," and the AR display highlights the location of the Settings app.
[1948] Output: Audio guide and augmented reality display presented to the user.
[1949] Step 7:
[1950] Process user feedback.
[1951] Input: The user interacts with the device through instructions and uses voice to ask questions or provide feedback, for example, "What should I do next?"
[1952] What it does: The device captures the user's voice feedback.
[1953] Output: Captured audio feedback data.
[1954] Step 8:
[1955] Performs voice recognition and emotion assessment.
[1956] Input: Captured audio feedback data.
[1957] Operation: The device uses a voice recognition means to convert voice data into text, and an emotion assessment means analyzes the user's voice tone and facial expressions to assess their emotions.
[1958] Output: Feedback and sentiment evaluation results converted into text.
[1959] Step 9:
[1960] Send the feedback to a cloud server.
[1961] Input: Feedback and sentiment rating results converted to text.
[1962] Operation: The device sends text data and emotion information to the cloud server using a communication method.
[1963] Output: The cloud server receives the feedback data and emotion information.
[1964] Step 10:
[1965] The cloud server generates a response.
[1966] Input: Feedback data and emotional information.
[1967] How it works: The server uses a generative AI model to generate appropriate responses based on user feedback, and flexibly adjusts the content and presentation of the response based on emotional assessment, for example, generating a calmer tone of response if the user is annoyed.
[1968] Output: The generated response text.
[1969] Step 11:
[1970] The terminal provides a response.
[1971] Input: The generated response text.
[1972] How it works: The device converts the response text into an audio file and augmented reality display data, and presents them to the user. Specifically, the audio prompt says "Select Wi-Fi settings" and the corresponding part is highlighted in the AR display.
[1973] Output: Audio guide and augmented reality display presented to the user.
[1974] Through the above processing steps, this system can provide appropriate and flexible support for operating digital devices while taking into account the user's emotions.
[1975] (Application example 2)
[1976] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1977] Elderly users face challenges when operating complex digital devices or selecting products in physical stores, such as difficulty obtaining and understanding information about operation and purchasing. Additionally, elderly users often experience anxiety and frustration while operating devices, which can be difficult to deal with appropriately. This can make it difficult for elderly users to smoothly navigate digital devices and physical stores.
[1978] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1979] In this invention, the server includes a camera, a communication unit, an analysis unit using a generative AI model, an emotion evaluation unit, a guide unit, a dialogue unit, and a voice recognition unit. This allows elderly users to easily obtain and understand product information on digital devices and in physical stores using AR glasses, and also allows them to receive flexible operation guidance that takes the user's emotions into consideration.
[1980] "AR glasses" are devices worn by elderly users that can overlay digital information onto their field of vision.
[1981] "Camera means" refers to a device for capturing an image of the user interface of a device within the user's field of view.
[1982] "Communication means" is a function for transmitting captured image data to a cloud server.
[1983] A "generative AI model" is an artificial intelligence model that analyzes received image data and generates appropriate operating procedures.
[1984] "Analysis means" is a function for analyzing transmitted images using a generative AI model.
[1985] The "guide means" is a function that provides users with operating procedures generated based on the analysis results through audio guidance and AR displays.
[1986] The "emotion evaluation means" is a function for analyzing the user's voice and facial expressions and evaluating the user's emotional state.
[1987] "Interactive means" is a function that accepts voice feedback from the user and provides additional operational guidance and emotion-sensitive guidance based on that feedback.
[1988] "Voice recognition means" is a function for converting the user's voice feedback into text data.
[1989] This invention is a system that makes it easier for elderly users to select products using complex digital devices or in physical stores. The system includes AR glasses worn by the elderly user, a camera means, a communication means, an analysis means having a generative AI model, an emotion evaluation means, a guidance means, a dialogue means, and a voice recognition means.
[1990] System Configuration
[1991] 1. AR Glasses:
[1992] This device is worn by the user and has a built-in camera, microphone, and speaker, allowing digital information to be displayed overlaid on the user's field of vision when operating a digital device or selecting products in a physical store.
[1993] 2. Camera means:
[1994] It is built into AR glasses and captures images of the device's user interface and products in physical stores within the user's field of view.
[1995] 3. Means of communication:
[1996] The captured image and audio data are sent to a cloud server using a secure, low-latency protocol.
[1997] 4. Analytical means with generative AI models:
[1998] The image data received by the server is analyzed and appropriate operating procedures and product information are generated. The generative AI model learns from a large data set and provides highly accurate analysis results.
[1999] 5. Emotional assessment measures:
[2000] It analyzes the user's tone of voice, facial expressions, and posture to assess the user's emotional state, allowing it to process user feedback appropriately and provide emotionally sensitive assistance.
[2001] 6. Guide means:
[2002] Based on the analysis results, the system generates specific next steps and product information, and provides these instructions to the user in the form of audio guidance or AR displays.
[2003] 7. Means of interaction:
[2004] If the user is unsure of a procedure or needs additional information, the system accepts voice feedback and provides additional operating instructions or emotionally sensitive guidance based on the feedback.
[2005] 8. Voice Recognition Methods:
[2006] The user's voice feedback is converted into text data, which is further analyzed by generative AI models and sentiment assessment methods.
[2007] Example
[2008] Example 1: Smartphone settings
[2009] The user puts on the AR glasses and projects the smartphone's home screen onto the camera. The cloud server analyzes this image using a generative AI model and provides the instruction, "Tap the Settings app." The AR glasses present this instruction to the user through voice guidance and AR displays. If the user is unsure of the next step, they can ask aloud, "What should I do next?" If the emotion engine detects the user's frustration along with this question, the server generates and provides more detailed instructions in a gentler tone.
[2010] Example 2: Product selection in a store
[2011] A user enters a supermarket and wants to look at a particular shelf. The AR glasses send an image of the shelf to a cloud server. The cloud server analyzes the image using a generative AI model and provides a voice prompt saying, "This product is rich in probiotics." If the user's emotions are unstable, the emotion assessment mechanism detects this and provides instructions in an encouraging tone, such as, "Don't worry, check out this product next."
[2012] Prompt example
[2013] "What should I do next?"
[2014] "Is this the right button to press?"
[2015] As described above, this system allows elderly users to easily operate digital devices and select products in physical stores. Furthermore, the emotional support allows them to use the system with peace of mind.
[2016] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2017] Step 1:
[2018] A user wears AR glasses and places a digital device or product in their field of view. The camera built into the AR glasses captures an image of the user interface or product. This image data becomes the input for the system. The camera means performs real-time image capture.
[2019] Step 2:
[2020] The device sends image data captured by the camera to the cloud server via a communication means. At this time, the image data is compressed, encoded, and sent with low latency. Data transmission is performed safely and quickly using a communication protocol.
[2021] Step 3:
[2022] The server analyzes the received image data using a generative AI model. This analysis method identifies the user interface and product attributes in the image and generates operation procedures and product information as the analysis results. The input is image data, and the output is the analysis results.
[2023] Step 4:
[2024] Based on the analysis results, the server generates the next operation procedure and product information, which the guide means provides to the terminal in the form of audio guide and AR display. The guide means converts the generated text into an audio file and AR display data.
[2025] Step 5:
[2026] The user operates the device and checks the product according to the presented operating instructions and product information. If the user is unsure of the next step, voice feedback is provided through the device's microphone. A prompt such as "What should I do next?" is input.
[2027] Step 6:
[2028] The device captures the user's voice feedback and converts it into text data using a voice recognition means. At the same time, the emotion assessment means analyzes the user's voice tone and facial expressions to evaluate the user's emotional state. The input is voice feedback, and the output is text data and emotion information.
[2029] Step 7:
[2030] The text data and emotional information converted by the speech recognition means are sent to a cloud server. The server then inputs this data back into the generative AI model to generate an appropriate response. The content and presentation of the response are flexibly adjusted based on the analysis results of the emotional evaluation means.
[2031] Step 8:
[2032] The server converts the generated response into an audio file and AR display data and sends them to the terminal. Based on this, the guide means provides the user with a response again in the form of audio guidance and AR display. For example, an instruction such as "Please select Wi-Fi settings" is generated.
[2033] This series of steps allows elderly users to smoothly operate digital devices and select products in physical stores.
[2034] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2035] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2036] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2037] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2038] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2039] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2040] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2041] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2042] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2043] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2044] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2045] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2046] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2047] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2048] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2049] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2050] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2051] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2052] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2053] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2054] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2055] The following is further disclosed regarding the above embodiment.
[2056] (Claim 1)
[2057] a camera means for capturing an image of a user interface of a device operated by an elderly user using AR glasses that can be worn by the elderly user;
[2058] A communication means for transmitting the captured image to a cloud server;
[2059] analysis means having a generative artificial intelligence model for analyzing the transmitted image;
[2060] a guide means for generating an operation procedure based on the analysis result and providing the operation procedure to the user as an audio guide and an AR display;
[2061] A system that includes an interaction means for accepting voice feedback from a user regarding the operation of a device in use and providing additional operation guidance based on the feedback.
[2062] (Claim 2)
[2063] 10. The system of claim 1, wherein the operating device is a smartphone, a tablet, a consumer electronics device, or a computer.
[2064] (Claim 3)
[2065] 10. The system of claim 1, further comprising a speech recognition means for converting the user's voice feedback into text using speech recognition technology.
[2066] "Example 1"
[2067] (Claim 1)
[2068] An imaging means for capturing an image of a user interface of an electronic device operated by an elderly user using AR glasses that can be worn by the elderly user;
[2069] communication means for transmitting the captured images to a remote server;
[2070] analysis means having a generative artificial intelligence model for analyzing the transmitted image;
[2071] an operation guidance means for generating an operation procedure based on the analysis result and providing the operation procedure to the user as an audio guide and an AR display;
[2072] an interaction means for receiving voice feedback from a user regarding the operation of the electronic device being used and providing additional operation guidance based on the feedback;
[2073] a speech recognition means for converting the user's voice feedback into text using speech recognition technology;
[2074] a presentation means for automatically presenting to a user an operation instruction generated based on the analyzed data;
[2075] A system including:
[2076] (Claim 2)
[2077] 2. The system of claim 1, wherein the operating electronic device is a mobile communication terminal, a tablet terminal, a home electronic device, or a calculator.
[2078] (Claim 3)
[2079] 2. The system according to claim 1, further comprising a generating means for generating a prompt sentence which is an input sentence to be given to the generative artificial intelligence model.
[2080] "Application Example 1"
[2081] (Claim 1)
[2082] An imaging means for capturing an image of a user interface of a device operated by an elderly user using AR glasses that can be worn by the elderly user;
[2083] a data communication means for transmitting the captured image to a cloud computing server;
[2084] analysis means having a generative artificial intelligence model for analyzing the transmitted image;
[2085] an operation guide means for generating an operation procedure based on the analysis result and providing the operation procedure to the user as an audio guide and an augmented reality display;
[2086] A system including an interaction means for receiving voice feedback from a user regarding the operation of a device in use and providing additional operation guidance based on the feedback,
[2087] Furthermore, an electronic payment support means that guides elderly users through transaction procedures when making electronic payments using their smartphones;
[2088] It includes a transaction guide means that instructs users to scan the QR code or tap the payment button depending on the progress of the transaction.
[2089] A system including:
[2090] (Claim 2)
[2091] 10. The system of claim 1, wherein the operating device is a smartphone, a tablet, a consumer electronics device, or a computer.
[2092] (Claim 3)
[2093] 10. The system of claim 1, further comprising a speech recognition means for converting the user's voice feedback into text using speech recognition technology.
[2094] "Example 2: Combining Emotion Engines"
[2095] (Claim 1)
[2096] An imaging means for capturing an image of a user interface of a device operated by an elderly user using AR glasses that can be worn by the elderly user;
[2097] a communication means for transmitting the captured image and audio data to a cloud server;
[2098] analysis means having a generative artificial intelligence model for analyzing the transmitted image and audio data;
[2099] a guide means for generating an operation procedure based on the analysis result and providing the operation procedure to the user as an audio guide and an augmented reality display;
[2100] an interaction means for receiving user voice feedback and providing additional operation guidance based on the feedback;
[2101] a speech recognition means for converting the user's voice feedback into text using speech recognition technology;
[2102] The system includes an emotion assessment means that assesses a user's emotions by analyzing voice tone, facial expressions, and posture.
[2103] (Claim 2)
[2104] 10. The system of claim 1, wherein the operating device is a smartphone, a tablet, a consumer electronics device, or a computer.
[2105] (Claim 3)
[2106] The system according to claim 1, characterized in that when the voice recognition means converts the user's voice feedback into text data, the emotion evaluation means analyzes the user's emotions and flexibly adjusts the generated operation guide.
[2107] "Application example 2 when combining emotion engines"
[2108] (Claim 1)
[2109] a camera means for capturing an image of a user interface of a device operated by an elderly user using AR glasses that can be worn by the elderly user;
[2110] A communication means for transmitting the captured image to a cloud server;
[2111] analysis means having a generative AI model for analyzing the transmitted image;
[2112] a guide means for generating an operation procedure based on the analysis result and providing the operation procedure to the user as an audio guide and an AR display;
[2113] The system includes an emotion assessment means for analyzing the user's voice feedback and emotional state regarding the operation of the device in use, and includes an interaction means for providing additional operation guides and emotion-sensitive guides based on the feedback.
[2114] (Claim 2)
[2115] 2. The system of claim 1, wherein the operation device is a smartphone, a tablet, a household appliance, a product installed in a store, or a computer.
[2116] (Claim 3)
[2117] 10. The system of claim 1, further comprising: converting the user's voice feedback into text using speech recognition technology; and analyzing the user's emotional state using an emotion assessment means. [Explanation of symbols]
[2118] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a camera means for capturing an image of a user interface of a device operated by an elderly user using AR glasses that can be worn by the elderly user; A communication means for transmitting the captured image to a cloud server; analysis means having a generative artificial intelligence model for analyzing the transmitted image; a guide means for generating an operation procedure based on the analysis result and providing the operation procedure to the user as an audio guide and an AR display; A system that includes an interaction means for accepting voice feedback from a user regarding the operation of a device in use and providing additional operation guidance based on the feedback.
2. The system of claim 1 , wherein the operating device is a smartphone, a tablet, a consumer electronics device, or a computer.
3. 10. The system of claim 1, further comprising a speech recognition means for converting the user's voice feedback into text using speech recognition techniques.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A