system

A system on the user's terminal uses OCR and computer vision to analyze visual information and provide solutions, addressing the complexity of smartphone operations and reducing stress for elderly users.

JP2026069136APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Elderly individuals face difficulties in using smartphones due to the complexity of operations and distinguishing between advertisements and important notifications, leading to stress and avoidance of digital devices.

Method used

A system installed on the user's terminal that analyzes visual information using OCR and computer vision to identify problems, providing solutions through visual and auditory means, ensuring an intuitive interface for easy device usage.

Benefits of technology

Enables elderly users to operate smartphones with confidence by quickly resolving issues and reducing stress through real-time problem identification and solution presentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069136000001_ABST
    Figure 2026069136000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] An analysis means installed on the user's terminal and used to analyze visual information displayed on the screen, A means for identifying problems faced by users using data obtained by the said analysis means, A means of presenting solutions to the user based on the problem identified by the specified means, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] For many modern elderly people, operating a smartphone is complex, and it is particularly difficult to distinguish between advertisements and important notifications, resulting in difficulties in daily communication and information management. As a result, there is a problem that they feel stressed and tend to avoid using digital devices. Therefore, there is a need to provide an intuitive and simple interface to enable the elderly to use smartphones with confidence.

Means for Solving the Problems

[0005] This invention solves these problems by providing a system that is installed on the user's operating terminal. This system includes multiple means for analyzing visual information on the screen in real time and identifying problems faced by the user. The analysis means uses OCR, computer vision, etc., to precisely recognize the visual information and extract the information the user needs. Furthermore, it provides appropriate support to the user by presenting solutions based on the identified problems. This presentation means provides the user with the necessary information in both visual and auditory senses, and through technological intervention, realizes an environment in which elderly people can use devices with peace of mind.

[0006] A "user" is an individual who uses the system, including elderly people, and who operates devices such as smartphones.

[0007] An "operating terminal" refers to a device that the user directly operates, such as a smartphone or tablet.

[0008] "Visual information" refers to all visual data, including text, images, and graphics, that are displayed on the screen of a device.

[0009] "Analysis means" refers to devices or software functions that analyze visual information, and which process information using OCR or computer vision technology.

[0010] "Identification means" refers to the process or function of recognizing and identifying problems faced by users using data obtained through analytical means.

[0011] "Presentation methods" refer to functions that convey information or solutions to users through visual and auditory means, and function as part of the user interface.

[0012] A "solution" refers to the methods or actions presented to resolve a problem faced by a user.

[0013] "Monitoring means" refers to a device or process that observes the user's actions in real time and collects data.

[0014] "Speech synthesis technology" refers to technology that converts text information into speech and provides it to users in an easily understandable format. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is an AI assistant system that helps elderly people use smartphones more easily. This system is installed as software on the user's smartphone, which is the user's operating terminal, and analyzes visual information on the screen to identify problems the user is facing.

[0037] First, the device uses its built-in screen capture function to constantly monitor the screen content. The acquired visual information is analyzed as text using OCR (Optical Character Recognition) technology, and images and patterns are recognized through computer vision. This allows the device to classify different elements such as advertisements, alerts, or messages.

[0038] Next, the device monitors the user's touch input and operation logs, identifying which operations the user is having difficulty with based on delays and multiple attempts at specific operations. This information is sent to a server, which uses machine learning algorithms to analyze it and generate solutions needed by the user. These solutions are expressed in concise and clear language to aid the user's understanding.

[0039] The generated solution is displayed as a visual notification on the device or provided to the user as an audio message using speech synthesis technology. For example, if a user cannot find the delete button in an email app, the device will instruct them via audio or on-screen message, "The delete button is in the upper right corner." This embodiment enables the rapid and effective resolution of problems faced by users and provides an environment where even elderly people can easily use smartphones.

[0040] As a concrete example of this system, consider a scenario where an advertisement appears while a user is using a social media app. In this situation, the device detects the advertisement and notifies the user with a message such as, "This is an advertisement. No action is required," supporting the user so they can continue using the app without confusion. This allows users to use digital devices with greater peace of mind.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The device periodically captures the user's smartphone screen using its built-in screen capture function. This captured data is used as visual information for subsequent analysis.

[0044] Step 2:

[0045] The device analyzes the visual information of the captured screen using OCR (Optical Character Recognition) technology. This extracts the text information on the screen as text data. Furthermore, it utilizes computer vision to recognize non-text elements, including images and icons.

[0046] Step 3:

[0047] The device sends the acquired text and image data to the server. The server receives the data and applies machine learning algorithms to classify different elements such as advertisements, alerts, and messages. This classification creates a foundation for understanding what information the user is interacting with.

[0048] Step 4:

[0049] The device collects user touch events and interaction logs. If a user repeatedly taps a specific button or stays on the same screen for an extended period, it detects and logs this. This helps determine if the user may be experiencing difficulties.

[0050] Step 5:

[0051] The server identifies the problems the user is facing based on the operation logs and classification data received from the terminal. For example, it may identify cases where the user has difficulty distinguishing between advertisements and content, or where the user does not understand the operating procedure.

[0052] Step 6:

[0053] The server generates a solution to the problem and creates a message to provide to the user in a concise and easy-to-understand format. This message is output as text that also serves as audio guidance, along with visual instructions.

[0054] Step 7:

[0055] The terminal notifies the user of a solution message received from the server. This can be done by displaying a pop-up message on the screen or by providing voice instructions to the user using speech synthesis technology.

[0056] Step 8:

[0057] The user follows instructions from the terminal and attempts to resolve the problem. After resolving the issue, the terminal monitors the user's actions again and sends the results as feedback to the server. This feedback is used to improve the system and enhance its accuracy.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] The present invention aims to provide a technology that effectively solves the operational difficulties that arise when users, including the elderly, use operating terminals due to the large amount of information on the screen and complex operations. Conventional technologies have made it difficult to identify users' operational problems in real time and provide appropriate support immediately. As a result, this has often compromised the convenience and sense of security of users.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes an analysis means installed on the operating device for analyzing image information displayed on the display surface, a monitoring means for monitoring user input operations and operation history and identifying user difficulties from delays and multiple attempts, and a notification means for notifying the user of the generated solution using visual or speech synthesis technology. This enables the rapid identification of problems in the user's operation and the provision of appropriate solutions.

[0063] An "operating device" is an electronic terminal device that users can directly operate, and it serves as the foundation on which this system is installed.

[0064] "Analysis means" refers to means that have the function of analyzing image information displayed on a display surface and identifying text data and other patterns.

[0065] "Identification means" refers to a means of identifying the problem a user is facing using information obtained through analysis means.

[0066] A "presentation means" is a means that has the function of providing a solution to the user based on the problem identified by the identification means.

[0067] "Monitoring means" refers to methods for monitoring user input operations and operation history in real time to identify delays or errors in operations.

[0068] A "generation method" is a means of sending information collected by a monitoring method to a server and using machine learning to create the optimal solution for the user.

[0069] A "notification method" is a means of communicating the generated solution to the user, and it uses visual displays or speech synthesis technology to notify the information.

[0070] This invention relates to an assistance system installed on electronic terminals operated by users, particularly the elderly. Embodiments of this system are described below.

[0071] A terminal refers to an electronic device with a user interface, such as a smartphone or tablet, where a program is executed. The terminal has a built-in screen capture function, which periodically captures information from the display screen. The captured images are analyzed using OCR (Optical Character Recognition) software and computer vision technology to identify text data and other visual information. This allows, for example, determining how advertisements or notification messages are displayed on the screen.

[0072] The terminal also features transparent monitoring capabilities, collecting user touch input and operation logs. This data provides crucial clues for analyzing where users are experiencing difficulties with operation. For example, if a particular button is frequently pressed incorrectly, the terminal records that operation and requests an appropriate solution from the server.

[0073] The server receives information sent from the terminal and analyzes the problem using machine learning algorithms. It then uses a generative AI model to generate solutions tailored to individual users. These solutions are provided to the user on the terminal via visual notifications or voice messages using speech synthesis technology. For example, clear instructions such as "The delete button is in the upper right corner" may be given.

[0074] For example, if an advertisement appears while a user is using a social media app, the device recognizes the advertisement and notifies the user with a message saying, "This is an advertisement. No action is required." In this way, users can avoid confusion caused by unnecessary actions and continue to use electronic devices with peace of mind.

[0075] As an example of a prompt, the model might be given an instruction such as, "Please suggest solutions to make smartphones easier for the elderly to use," and appropriate support content will be generated accordingly.

[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0077] Step 1:

[0078] The device periodically captures screen content using its built-in screen capture function. The input consists of all screen information displayed on the device's screen. The captured images are then directly sent as output for OCR and computer vision analysis. This process is performed every few seconds, allowing the user to evaluate the applications and notifications they are currently using.

[0079] Step 2:

[0080] The device converts the acquired screenshot into text data using OCR technology and then uses computer vision to identify images and patterns. The input is a screenshot image, and the output is text data and associated pattern information. Based on this information, it identifies elements such as advertisements and icons within apps. As a specific example, it identifies advertising banners displayed on social media apps.

[0081] Step 3:

[0082] The device monitors the user's touch input and operation history, and records the information. Input includes the user's screen touch location and time, and output is a log of that operation history. If repeated errors occur over a certain period, the data is used to identify difficult operations. Specifically, this involves counting the number of times the same icon is tapped repeatedly.

[0083] Step 4:

[0084] The device sends information about the identified user's operational problems to the server. The input is data about the identified operation history, and the output is a request to the server. This is sent to the server for further analysis. Specifically, it sends data on the number of times the user has made a mistake deleting operation in the email application.

[0085] Step 5:

[0086] The server analyzes the received data using machine learning algorithms. The input is operation data sent from the terminal, and the output is a solution based on that data. A generative AI model is used to determine specific assistance. For example, it might generate guidance such as, "The delete button is in the upper right corner of the screen."

[0087] Step 6:

[0088] The device notifies the user of the solution received from the server. The input is the solution data from the server, and the output is the notification to the user. The notification is either displayed visually on the screen or played back using speech synthesis technology. For example, the user is informed via voice, "You don't need to worry about the advertisements."

[0089] (Application Example 1)

[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0091] When elderly people use digital devices, especially when dealing with complex electronic payment systems, errors or inaccurate judgments can lead to problems. These challenges pose a barrier to the safe and secure use of electronic payments. Therefore, there is a need to provide an environment that allows users to make electronic payments more intuitively and securely.

[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0093] In this invention, the server includes analysis means for analyzing visual information installed on the operating device, identification means for identifying a problem faced by the user using the obtained information, notification means for presenting a solution to the user based on the identified problem, and support means for monitoring transaction screens and guiding the user to points where decision-making is required. This enables the user to proceed with complex electronic payment operations with confidence by following the guidance.

[0094] An "operating device" is an electronic device used by the user to perform operations, and in this context, it mainly refers to a smartphone.

[0095] "Visual information" refers to visual data such as text and images displayed on a screen.

[0096] "Analysis means" refers to technical means used to read visual information and understand its content.

[0097] "Information" refers to data and insights obtained through analytical methods.

[0098] An "identification method" is a function that identifies the challenges and problems that the user faces based on the information obtained.

[0099] "Notification method" refers to a method of informing the user of solutions to identified problems.

[0100] "Support measures" refer to functions that monitor transaction-related screens and guide users to points where decision-making is required.

[0101] "Transaction-related screens" refer to interface screens where processes such as electronic payments are processed.

[0102] "Decision-making" refers to the judgments that users need to make when proceeding with transactions or operations.

[0103] The system for realizing this invention is installed on the user's operating device and consists mainly of a program that processes various types of information. The server constantly monitors the visual information on the screen through analysis means running on the operating device and recognizes it as text using OCR (optical character recognition) technology. Furthermore, it uses computer vision to classify images and patterns and identify screens related to transactions.

[0104] The analysis means utilizes this visual information to help the identification means identify the user's problem. Based on the identified problem, the notification means uses speech synthesis technology to notify the user of appropriate actions or instructions from the server. For example, it provides clear guidance to the user regarding the location of the button for completing a transaction and the steps involved in its operation.

[0105] Furthermore, the support system monitors the transaction selection screen when the user is making an electronic payment, providing voice and on-screen instructions at appropriate points where the user needs to make a decision. This allows the user to continue the operation safely.

[0106] As a concrete example, consider a scenario where an elderly user is making an online purchase. In this case, a voice prompt such as, "Tap this button to complete the payment," would be provided, allowing the user to complete the purchase without confusion.

[0107] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0108] Step 1:

[0109] The device periodically acquires visual information from the screen using its screen capture function. The input is the visual information displayed on the device's screen, and the output is the storage of this information as image data. This acquired data is then converted into text data using OCR technology.

[0110] Step 2:

[0111] The server analyzes the acquired text data and uses identification methods to identify potential problems the user may be facing. The input is text data converted by OCR, and the output is a clarification of the issues the user needs to address. It analyzes words and phrases within the text to determine which information is important.

[0112] Step 3:

[0113] The device presents appropriate solutions to identified problems. It uses notification methods and speech synthesis technology to present solutions to the user. The input is the clearly defined problem the user faces, and the output is voice guidance or on-screen instructions. For example, it provides specific voice instructions such as, "Tap this button to complete the payment."

[0114] Step 4:

[0115] The support system monitors transaction-related screens and assists users in performing transactions. Input is the state of the transaction screen, and output is user guidance regarding necessary actions. It recognizes specific elements of the transaction screen and generates messages prompting the user to take the next action.

[0116] Step 5:

[0117] The user makes decisions by following the instructions provided by the device. The input is the instructional messages from the device, and the output is the user's choice of action. Based on the information provided, the user can proceed with the necessary actions without hesitation.

[0118] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0119] This invention is an AI assistant system that helps elderly people use smartphones more effectively, and in particular, incorporates a function to recognize the user's emotions. The system is installed on the user's terminal, analyzes various data inputs in real time, identifies the problems they face, and then presents solutions that are tailored to their emotions.

[0120] First, the device continuously collects visual information on the screen using its screen capture function. This data is analyzed using OCR technology, recognizing image elements through computer vision. This allows the information on the screen to be analyzed and classified into elements such as advertisements, alerts, and general messages.

[0121] Next, the device senses the user's voice and facial expressions and analyzes those emotions using its built-in emotion engine. By combining data such as voice tone, changes in facial expressions, and even operation speed, it infers the user's emotional state, such as whether they are feeling stressed or comfortable using the device.

[0122] This information is sent to a server, which uses machine learning algorithms to analyze the user's current situation and emotional state. Based on these results, the server identifies the specific problem the user is experiencing and generates an appropriate solution. The solution is tailored to the user's emotions and is designed to provide reassurance.

[0123] The generated solutions are sent to the device and notified to the user. These notifications include both on-screen visuals and voice guidance utilizing speech synthesis technology. For example, if a user becomes confused and frustrated while using the app, the system provides an encouraging message such as, "Try this, you'll get the hang of it." In this way, the system supports older adults in using technology with confidence through user-emotion-based feedback.

[0124] The following describes the processing flow.

[0125] Step 1:

[0126] The device uses a screen capture function to periodically capture the user's smartphone screen and collect data as visual information.

[0127] Step 2:

[0128] The device uses OCR technology to extract text from captured visual information, and further recognizes and classifies images and icons using computer vision.

[0129] Step 3:

[0130] The device uses a camera and microphone to detect the user's voice and facial expressions, and an emotion engine analyzes their emotional state. This analysis determines the user's stress levels, happiness, and other emotional states.

[0131] Step 4:

[0132] The terminal sends these analysis results to the server, which uses the received data to identify the problems the user is facing using machine learning algorithms.

[0133] Step 5:

[0134] The server generates customized solutions to resolve problems based on the user's emotional state. These solutions are designed as messages that are considerate and reassuring to the user.

[0135] Step 6:

[0136] The terminal receives the solution sent from the server and notifies the user. This notification is provided both as a visual message on the UI and as voice guidance using speech synthesis technology.

[0137] Step 7:

[0138] The user attempts to resolve the problem by following the instructions provided by the terminal. The terminal collects the results of these operations as feedback and sends it to the server. This feedback is used to improve the system.

[0139] (Example 2)

[0140] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0141] In modern society, elderly people face difficulties in effectively using smartphones. The complexity of operation and information overload can cause stress and lead to overlooking important information. Therefore, there is a need for systems that provide operational support and emotionally sensitive responses to enable elderly people to use smartphones with peace of mind.

[0142] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0143] In this invention, the server includes a collection means for continuously collecting the user's visual information and analyzing textual information, a classification means for classifying the visual information, and an emotion recognition means for analyzing the emotional state. This makes it possible to identify the problems the user is facing and provide solutions that give a sense of security in line with their emotions.

[0144] "Collection means" refers to a function on the user's terminal that continuously collects visual information displayed on the screen and analyzes the textual information using optical character recognition technology.

[0145] A "classification method" is a function that uses computer vision technology to analyze visual information acquired by collection methods and identify and classify elements such as advertisements, alerts, and general messages.

[0146] "Emotion recognition means" refers to a function that senses the user's voice and facial expressions, and infers the user's emotional state by analyzing voice tone, facial expression changes, and the speed of a series of movements.

[0147] "Identification means" refers to a function that uses data obtained from classification means and emotion recognition means to identify specific problems and challenges that the user is facing.

[0148] The "generation method" refers to a function that applies machine learning algorithms to data sent to the server to generate appropriate solutions. The solutions are adjusted according to the user's emotional state.

[0149] "Presentation means" refers to a function that provides users with solutions created by the generation means using visual display or speech synthesis technology.

[0150] This invention is a support system for elderly people to use smartphones more effectively, providing an AI assistant combined with emotion recognition capabilities. This system is installed on the user's terminal and utilizes hardware such as screen capture functionality, microphones, and cameras built into the terminal, as well as software such as OCR technology and computer vision technology.

[0151] The device collects information on the screen in real time using its screen capture function and analyzes that information using OCR technology. The data extracted through this analysis is then classified into elements such as advertisements, alerts, and general messages on the screen using computer vision technology. In addition, the device senses the user's voice and facial expressions using a microphone and camera, and analyzes the voice tone and facial expressions to infer the user's emotional state.

[0152] These analysis results are sent to a server, which uses machine learning algorithms to analyze the user's current situation and emotional state. Based on this analysis, the server generates an appropriate solution and adjusts it according to the user's emotions. For example, if the user is feeling anxious, the server will generate a message using reassuring language. The generated solution is sent to the terminal and notified to the user. The terminal can not only display the solution visually on the screen but can also provide voice guidance using speech synthesis technology.

[0153] Specifically, when a user gets lost in the app's settings, they will be given specific voice guidance such as, "Tap the gear icon here, and then select 'Network Settings'." In this way, users can instantly receive feedback that matches their emotional state at that moment.

[0154] An example of a prompt might be, "If a user is struggling to understand a new feature in the app, how can we guide them to a solution while also providing reassurance?" This prompt allows the generative AI model to derive and present individual solutions tailored to the user's situation.

[0155] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0156] Step 1:

[0157] The device collects visual information on the screen in real time using its screen capture function. This collected data is then analyzed using OCR technology to extract text data. The input is the visual information on the screen, and the output is the analyzed text data. In this process, the device prepares the basic data necessary to identify elements such as advertisements, alerts, and general messages.

[0158] Step 2:

[0159] The device uses computer vision technology to analyze image elements on the screen and combine them with collected text data to classify advertisements, alerts, general messages, etc. The input is text data and image data extracted by OCR, and the output is classified visual information. This classification process clarifies the type of information the user is facing on the screen.

[0160] Step 3:

[0161] The device uses a microphone and camera to sense the user's voice and facial expressions. This allows for the collection of voice tone and facial expression data. The input is the user's real-time voice and facial expressions captured by the microphone and camera, and the output is data for emotion estimation. The device uses this data to infer the user's emotional state.

[0162] Step 4:

[0163] The device sends collected visual and emotional data to the server. The server analyzes this data using machine learning algorithms. The input is the classified visual and emotional data sent from the device, and the output is the analyzed user situation and emotional state. Based on this analysis, the server identifies the problems the user is facing.

[0164] Step 5:

[0165] Based on the identified problem, the server uses a generative AI model to generate appropriate solutions that are sensitive to the user's emotions. The input is the server's analysis of the user's situation and problem, and the output is a customized solution. The generated solution is adjusted according to the user's emotional state.

[0166] Step 6:

[0167] The server sends the generated solution to the terminal. The terminal receives this solution and notifies the user. The notification is provided both visually and through voice guidance using speech synthesis technology. The input is the solution sent from the server, and the output is the visual and voice notification to the user. This allows the user to receive specific instructions and confidently take the next action.

[0168] (Application Example 2)

[0169] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0170] When elderly people use communication devices in commercial facilities and other locations, they often find them difficult to operate due to their unfamiliarity with the technology, which can increase stress and anxiety. This can also impair their user experience and reduce convenience. There is a need to provide systems that allow elderly people to operate communication devices with greater confidence and enjoy a more fulfilling user experience.

[0171] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0172] In this invention, the server includes means for analyzing visual displays installed on the user's communication device, means for identifying the problem the user is facing, means for presenting solutions, and means for recognizing the user's emotional state and adjusting the solutions. This makes it possible for elderly people to use communication devices safely and without stress in commercial facilities and other places.

[0173] A "communication device" is a device used for voice communication, data communication, etc., and is a device that allows users to input and receive information.

[0174] "Analysis means" refers to a device or program that has the function of analyzing visual displays or input data and understanding or identifying their content.

[0175] "Identification means" refers to the technology or process used to identify and clarify the challenges faced by users, using analyzed data.

[0176] A "presentation method" is a system that provides users with appropriate solutions or information based on problems identified through specific means.

[0177] "Emotion recognition means" refers to technologies and methods that analyze a user's facial expressions, voice, and actions to identify and evaluate their psychological state and emotions.

[0178] The system that realizes this invention will enable elderly people to safely obtain information using communication devices in commercial facilities and other locations. The server will function via an application installed on the communication device and will utilize various hardware and software.

[0179] The communication device includes hardware such as cameras, microphones, speakers, and displays. The analysis method uses this hardware to acquire image data and analyze the visual display. The software used includes computer vision technology (e.g., OCR technology) and emotion recognition technology (e.g., machine learning algorithms). Specifically, for image analysis, services such as Amazon Rekognition and Google Cloud Vision can be applied, and for emotion recognition, methods that read emotions from voice and facial expressions are employed.

[0180] The server processes data collected by the analysis means and identifies the user's challenges using the identification means. Based on these identified challenges, the presentation means presents appropriate solutions and information to the user via voice or text display. The emotion recognition means understands the user's psychological state and adjusts the solutions according to the user's emotions. As a result, elderly people can operate communication devices smoothly and without stress when using them in commercial facilities.

[0181] As a concrete example, if a user is lost in the store, the camera on the communication device reads and analyzes the store's map information. If the user appears frustrated, a gentle voice message such as, "Hello! The special sale section is over here. Do you need any guidance?" and a simple map display are provided.

[0182] Examples of input prompts for a generative AI model are as follows:

[0183] "Please create a message that will appropriately guide elderly customers who are lost in the store. It should be especially reassuring for those who are unsure about how to use the machines."

[0184] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0185] Step 1:

[0186] The device uses its camera to acquire visual information about its surroundings. This input data is in image format and is transmitted to the analysis system. The device uses OCR technology to convert the image data into text data and recognize store maps and product labels.

[0187] Step 2:

[0188] The device uses identified text data and collected location information (e.g., GPS or in-store beacons) to determine the user's current location and what they are looking at. This data is sent to a server to identify the user's location and interests.

[0189] Step 3:

[0190] The server analyzes the data sent from the terminal and forms prompt sentences to input into the AI ​​model. These prompt sentences are adjusted based on the user's current location, interests, and past behavioral history. Furthermore, the server evaluates the user's challenges and stress levels based on data from emotion recognition systems.

[0191] Step 4:

[0192] The server requests a solution from the AI ​​model based on the prompt message and generates a guidance message as output. For example, it might create a message such as, "The special sale section is over here. Do you need guidance?"

[0193] Step 5:

[0194] The generated guidance message is sent to the terminal and visually displayed to the user through a display device, as well as provided as voice guidance using speech synthesis technology. This allows the user to obtain guidance information from their current location to their destination, supporting them in traveling without stress.

[0195] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0196] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0197] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0198] [Second Embodiment]

[0199] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0200] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0201] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0202] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0203] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0204] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0205] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0206] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0207] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0208] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0209] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0210] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0211] This invention is an AI assistant system that helps elderly people use smartphones more easily. This system is installed as software on the user's smartphone, which is the user's operating terminal, and analyzes visual information on the screen to identify problems the user is facing.

[0212] First, the device uses its built-in screen capture function to constantly monitor the screen content. The acquired visual information is analyzed as text using OCR (Optical Character Recognition) technology, and images and patterns are recognized through computer vision. This allows the device to classify different elements such as advertisements, alerts, or messages.

[0213] Next, the device monitors the user's touch input and operation logs, identifying which operations the user is having difficulty with based on delays and multiple attempts at specific operations. This information is sent to a server, which uses machine learning algorithms to analyze it and generate solutions needed by the user. These solutions are expressed in concise and clear language to aid the user's understanding.

[0214] The generated solution is displayed as a visual notification on the device or provided to the user as an audio message using speech synthesis technology. For example, if a user cannot find the delete button in an email app, the device will instruct them via audio or on-screen message, "The delete button is in the upper right corner." This embodiment enables the rapid and effective resolution of problems faced by users and provides an environment where even elderly people can easily use smartphones.

[0215] As a concrete example of this system, consider a scenario where an advertisement appears while a user is using a social media app. In this situation, the device detects the advertisement and notifies the user with a message such as, "This is an advertisement. No action is required," supporting the user so they can continue using the app without confusion. This allows users to use digital devices with greater peace of mind.

[0216] The following describes the processing flow.

[0217] Step 1:

[0218] The device periodically captures the user's smartphone screen using its built-in screen capture function. This captured data is used as visual information for subsequent analysis.

[0219] Step 2:

[0220] The device analyzes the visual information of the captured screen using OCR (Optical Character Recognition) technology. This extracts the text information on the screen as text data. Furthermore, it utilizes computer vision to recognize non-text elements, including images and icons.

[0221] Step 3:

[0222] The device sends the acquired text and image data to the server. The server receives the data and applies machine learning algorithms to classify different elements such as advertisements, alerts, and messages. This classification creates a foundation for understanding what information the user is interacting with.

[0223] Step 4:

[0224] The device collects user touch events and interaction logs. If a user repeatedly taps a specific button or stays on the same screen for an extended period, it detects and logs this. This helps determine if the user may be experiencing difficulties.

[0225] Step 5:

[0226] The server identifies the problems the user is facing based on the operation logs and classification data received from the terminal. For example, it may identify cases where the user has difficulty distinguishing between advertisements and content, or where the user does not understand the operating procedure.

[0227] Step 6:

[0228] The server generates a solution to the problem and creates a message to provide to the user in a concise and easy-to-understand format. This message is output as text that also serves as audio guidance, along with visual instructions.

[0229] Step 7:

[0230] The terminal notifies the user of a solution message received from the server. This can be done by displaying a pop-up message on the screen or by providing voice instructions to the user using speech synthesis technology.

[0231] Step 8:

[0232] The user follows instructions from the terminal and attempts to resolve the problem. After resolving the issue, the terminal monitors the user's actions again and sends the results as feedback to the server. This feedback is used to improve the system and enhance its accuracy.

[0233] (Example 1)

[0234] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0235] The present invention aims to provide a technology that effectively solves the operational difficulties that arise when users, including the elderly, use operating terminals due to the large amount of information on the screen and complex operations. Conventional technologies have made it difficult to identify users' operational problems in real time and provide appropriate support immediately. As a result, this has often compromised the convenience and sense of security of users.

[0236] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0237] In this invention, the server includes an analysis means installed on the operating device for analyzing image information displayed on the display surface, a monitoring means for monitoring user input operations and operation history and identifying user difficulties from delays and multiple attempts, and a notification means for notifying the user of the generated solution using visual or speech synthesis technology. This enables the rapid identification of problems in the user's operation and the provision of appropriate solutions.

[0238] An "operating device" is an electronic terminal device that users can directly operate, and it serves as the foundation on which this system is installed.

[0239] "Analysis means" refers to means that have the function of analyzing image information displayed on a display surface and identifying text data and other patterns.

[0240] "Identification means" refers to a means of identifying the problem a user is facing using information obtained through analysis means.

[0241] A "presentation means" is a means that has the function of providing a solution to the user based on the problem identified by the identification means.

[0242] "Monitoring means" refers to methods for monitoring user input operations and operation history in real time to identify delays or errors in operations.

[0243] A "generation method" is a means of sending information collected by a monitoring method to a server and using machine learning to create the optimal solution for the user.

[0244] A "notification method" is a means of communicating the generated solution to the user, and it uses visual displays or speech synthesis technology to notify the information.

[0245] This invention relates to an assistance system installed on electronic terminals operated by users, particularly the elderly. Embodiments of this system are described below.

[0246] A terminal refers to an electronic device with a user interface, such as a smartphone or tablet, where a program is executed. The terminal has a built-in screen capture function, which periodically captures information from the display screen. The captured images are analyzed using OCR (Optical Character Recognition) software and computer vision technology to identify text data and other visual information. This allows, for example, determining how advertisements or notification messages are displayed on the screen.

[0247] The terminal also features transparent monitoring capabilities, collecting user touch input and operation logs. This data provides crucial clues for analyzing where users are experiencing difficulties with operation. For example, if a particular button is frequently pressed incorrectly, the terminal records that operation and requests an appropriate solution from the server.

[0248] The server receives information sent from the terminal and analyzes the problem using machine learning algorithms. It then uses a generative AI model to generate solutions tailored to individual users. These solutions are provided to the user on the terminal via visual notifications or voice messages using speech synthesis technology. For example, clear instructions such as "The delete button is in the upper right corner" may be given.

[0249] For example, if an advertisement appears while a user is using a social media app, the device recognizes the advertisement and notifies the user with a message saying, "This is an advertisement. No action is required." In this way, users can avoid confusion caused by unnecessary actions and continue to use electronic devices with peace of mind.

[0250] As an example of a prompt, the model might be given an instruction such as, "Please suggest solutions to make smartphones easier for the elderly to use," and appropriate support content will be generated accordingly.

[0251] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0252] Step 1:

[0253] The device periodically captures screen content using its built-in screen capture function. The input consists of all screen information displayed on the device's screen. The captured images are then directly sent as output for OCR and computer vision analysis. This process is performed every few seconds, allowing the user to evaluate the applications and notifications they are currently using.

[0254] Step 2:

[0255] The device converts the acquired screenshot into text data using OCR technology and then uses computer vision to identify images and patterns. The input is a screenshot image, and the output is text data and associated pattern information. Based on this information, it identifies elements such as advertisements and icons within apps. As a specific example, it identifies advertising banners displayed on social media apps.

[0256] Step 3:

[0257] The device monitors the user's touch input and operation history, and records the information. Input includes the user's screen touch location and time, and output is a log of that operation history. If repeated errors occur over a certain period, the data is used to identify difficult operations. Specifically, this involves counting the number of times the same icon is tapped repeatedly.

[0258] Step 4:

[0259] The device sends information about the identified user's operational problems to the server. The input is data about the identified operation history, and the output is a request to the server. This is sent to the server for further analysis. Specifically, it sends data on the number of times the user has made a mistake deleting operation in the email application.

[0260] Step 5:

[0261] The server analyzes the received data using machine learning algorithms. The input is operation data sent from the terminal, and the output is a solution based on that data. A generative AI model is used to determine specific assistance. For example, it might generate guidance such as, "The delete button is in the upper right corner of the screen."

[0262] Step 6:

[0263] The device notifies the user of the solution received from the server. The input is the solution data from the server, and the output is the notification to the user. The notification is either displayed visually on the screen or played back using speech synthesis technology. For example, the user is informed via voice, "You don't need to worry about the advertisements."

[0264] (Application Example 1)

[0265] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0266] When elderly people use digital devices, especially when dealing with complex electronic payment systems, errors or inaccurate judgments can lead to problems. These challenges pose a barrier to the safe and secure use of electronic payments. Therefore, there is a need to provide an environment that allows users to make electronic payments more intuitively and securely.

[0267] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0268] In this invention, the server includes analysis means for analyzing visual information installed on the operating device, identification means for identifying a problem faced by the user using the obtained information, notification means for presenting a solution to the user based on the identified problem, and support means for monitoring transaction screens and guiding the user to points where decision-making is required. This enables the user to proceed with complex electronic payment operations with confidence by following the guidance.

[0269] An "operating device" is an electronic device used by the user to perform operations, and in this context, it mainly refers to a smartphone.

[0270] "Visual information" refers to visual data such as text and images displayed on a screen.

[0271] "Analysis means" refers to technical means used to read visual information and understand its content.

[0272] "Information" refers to data and insights obtained through analytical methods.

[0273] An "identification method" is a function that identifies the challenges and problems that the user faces based on the information obtained.

[0274] "Notification method" refers to a method of informing the user of solutions to identified problems.

[0275] "Support measures" refer to functions that monitor transaction-related screens and guide users to points where decision-making is required.

[0276] "Transaction-related screens" refer to interface screens where processes such as electronic payments are processed.

[0277] "Decision-making" refers to the judgments that users need to make when proceeding with transactions or operations.

[0278] The system for realizing this invention is installed on the user's operating device and consists mainly of a program that processes various types of information. The server constantly monitors the visual information on the screen through analysis means running on the operating device and recognizes it as text using OCR (optical character recognition) technology. Furthermore, it uses computer vision to classify images and patterns and identify screens related to transactions.

[0279] The analysis means utilizes this visual information to assist the identification means in specifying the user's issues. For the identified issues, the notification means notifies the user from the server with appropriate actions or instructions by means of speech synthesis technology. For example, clear guidance is provided to the user regarding the position of the button for transaction completion and the operation procedure.

[0280] Also, when the user conducts an electronic payment, the support means monitors the selection screen related to the transaction and provides appropriate voice or screen instructions at the points where the user needs to make decisions. Thereby, the user can continue the operation safely.

[0281] As a specific example, consider the case where an elderly user conducts an online purchase procedure. At this time, prompt sentences such as "Tapping this button will complete the payment." are instructed by voice, enabling the user to complete the purchase without confusion.

[0282] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0283] Step 1:

[0284] The terminal periodically acquires the visual information on the screen using the screen capture function. The input is the visual information displayed on the terminal's display, and the output is to hold this as image data. The acquired data is converted into text data using OCR technology.

[0285] Step 2:

[0286] The server analyzes the acquired text data and uses the identification means to identify the possible problems the user is facing. The input is the text data converted by OCR, and the output is the clarification of the issues the user should address. Analyze the words and phrases contained in the text to determine which information is important.

[0287] Step 3:

[0288] The device presents appropriate solutions to identified problems. It uses notification methods and speech synthesis technology to present solutions to the user. The input is the clearly defined problem the user faces, and the output is voice guidance or on-screen instructions. For example, it provides specific voice instructions such as, "Tap this button to complete the payment."

[0289] Step 4:

[0290] The support system monitors transaction-related screens and assists users in performing transactions. Input is the state of the transaction screen, and output is user guidance regarding necessary actions. It recognizes specific elements of the transaction screen and generates messages prompting the user to take the next action.

[0291] Step 5:

[0292] The user makes decisions by following the instructions provided by the device. The input is the instructional messages from the device, and the output is the user's choice of action. Based on the information provided, the user can proceed with the necessary actions without hesitation.

[0293] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0294] This invention is an AI assistant system that helps elderly people use smartphones more effectively, and in particular, incorporates a function to recognize the user's emotions. The system is installed on the user's terminal, analyzes various data inputs in real time, identifies the problems they face, and then presents solutions that are tailored to their emotions.

[0295] First, the device continuously collects visual information on the screen using its screen capture function. This data is analyzed using OCR technology, recognizing image elements through computer vision. This allows the information on the screen to be analyzed and classified into elements such as advertisements, alerts, and general messages.

[0296] Next, the device senses the user's voice and facial expressions and analyzes those emotions using its built-in emotion engine. By combining data such as voice tone, changes in facial expressions, and even operation speed, it infers the user's emotional state, such as whether they are feeling stressed or comfortable using the device.

[0297] This information is sent to a server, which uses machine learning algorithms to analyze the user's current situation and emotional state. Based on these results, the server identifies the specific problem the user is experiencing and generates an appropriate solution. The solution is tailored to the user's emotions and is designed to provide reassurance.

[0298] The generated solutions are sent to the device and notified to the user. These notifications include both on-screen visuals and voice guidance utilizing speech synthesis technology. For example, if a user becomes confused and frustrated while using the app, the system provides an encouraging message such as, "Try this, you'll get the hang of it." In this way, the system supports older adults in using technology with confidence through user-emotion-based feedback.

[0299] The following describes the processing flow.

[0300] Step 1:

[0301] The device uses a screen capture function to periodically capture the user's smartphone screen and collect data as visual information.

[0302] Step 2:

[0303] The terminal uses OCR technology to extract text from the captured visual information, and further uses computer vision to recognize images and icons, and classifies each of them.

[0304] Step 3:

[0305] The terminal uses a camera and a microphone to sense the user's voice and expression, and analyzes the emotional state by an emotion engine. Based on this analysis, it judges the stress, happiness, etc. that the user is feeling.

[0306] Step 4:

[0307] The terminal sends these analysis results to the server, and the server uses the received data to identify the problems the user is facing with a machine learning algorithm.

[0308] Step 5:

[0309] The server generates a customized solution for problem-solving based on the user's emotional state. The solution is designed as a message that includes considerations to make the user feel at ease.

[0310] Step 6:

[0311] The terminal receives the solution sent from the server and notifies the user. This notification is carried out in both the form of message display on the visual UI and voice guidance using voice synthesis technology.

[0312] Step 7:

[0313] The user performs operations according to the instructions provided by the terminal and attempts to solve the problem. The terminal collects the operation results as feedback and sends them to the server. The feedback can be used to improve the system.

[0314] (Example 2)

[0315] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0316] In modern society, elderly people face difficulties in effectively using smartphones. The complexity of operation and information overload can cause stress and lead to overlooking important information. Therefore, there is a need for systems that provide operational support and emotionally sensitive responses to enable elderly people to use smartphones with peace of mind.

[0317] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0318] In this invention, the server includes a collection means for continuously collecting the user's visual information and analyzing textual information, a classification means for classifying the visual information, and an emotion recognition means for analyzing the emotional state. This makes it possible to identify the problems the user is facing and provide solutions that give a sense of security in line with their emotions.

[0319] "Collection means" refers to a function on the user's terminal that continuously collects visual information displayed on the screen and analyzes the textual information using optical character recognition technology.

[0320] A "classification method" is a function that uses computer vision technology to analyze visual information acquired by collection methods and identify and classify elements such as advertisements, alerts, and general messages.

[0321] "Emotion recognition means" refers to a function that senses the user's voice and facial expressions, and infers the user's emotional state by analyzing voice tone, facial expression changes, and the speed of a series of movements.

[0322] "Identification means" refers to a function that uses data obtained from classification means and emotion recognition means to identify specific problems and challenges that the user is facing.

[0323] The "generation method" refers to a function that applies machine learning algorithms to data sent to the server to generate appropriate solutions. The solutions are adjusted according to the user's emotional state.

[0324] "Presentation means" refers to a function that provides users with solutions created by the generation means using visual display or speech synthesis technology.

[0325] This invention is a support system for elderly people to use smartphones more effectively, providing an AI assistant combined with emotion recognition capabilities. This system is installed on the user's terminal and utilizes hardware such as screen capture functionality, microphones, and cameras built into the terminal, as well as software such as OCR technology and computer vision technology.

[0326] The device collects information on the screen in real time using its screen capture function and analyzes that information using OCR technology. The data extracted through this analysis is then classified into elements such as advertisements, alerts, and general messages on the screen using computer vision technology. In addition, the device senses the user's voice and facial expressions using a microphone and camera, and analyzes the voice tone and facial expressions to infer the user's emotional state.

[0327] These analysis results are sent to a server, which uses machine learning algorithms to analyze the user's current situation and emotional state. Based on this analysis, the server generates an appropriate solution and adjusts it according to the user's emotions. For example, if the user is feeling anxious, the server will generate a message using reassuring language. The generated solution is sent to the terminal and notified to the user. The terminal can not only display the solution visually on the screen but can also provide voice guidance using speech synthesis technology.

[0328] Specifically, when a user gets lost in the app's settings, they will be given specific voice guidance such as, "Tap the gear icon here, and then select 'Network Settings'." In this way, users can instantly receive feedback that matches their emotional state at that moment.

[0329] An example of a prompt might be, "If a user is struggling to understand a new feature in the app, how can we guide them to a solution while also providing reassurance?" This prompt allows the generative AI model to derive and present individual solutions tailored to the user's situation.

[0330] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0331] Step 1:

[0332] The device collects visual information on the screen in real time using its screen capture function. This collected data is then analyzed using OCR technology to extract text data. The input is the visual information on the screen, and the output is the analyzed text data. In this process, the device prepares the basic data necessary to identify elements such as advertisements, alerts, and general messages.

[0333] Step 2:

[0334] The device uses computer vision technology to analyze image elements on the screen and combine them with collected text data to classify advertisements, alerts, general messages, etc. The input is text data and image data extracted by OCR, and the output is classified visual information. This classification process clarifies the type of information the user is facing on the screen.

[0335] Step 3:

[0336] The device uses a microphone and camera to sense the user's voice and facial expressions. This allows for the collection of voice tone and facial expression data. The input is the user's real-time voice and facial expressions captured by the microphone and camera, and the output is data for emotion estimation. The device uses this data to infer the user's emotional state.

[0337] Step 4:

[0338] The device sends collected visual and emotional data to the server. The server analyzes this data using machine learning algorithms. The input is the classified visual and emotional data sent from the device, and the output is the analyzed user situation and emotional state. Based on this analysis, the server identifies the problems the user is facing.

[0339] Step 5:

[0340] Based on the identified problem, the server uses a generative AI model to generate appropriate solutions that are sensitive to the user's emotions. The input is the server's analysis of the user's situation and problem, and the output is a customized solution. The generated solution is adjusted according to the user's emotional state.

[0341] Step 6:

[0342] The server sends the generated solution to the terminal. The terminal receives this solution and notifies the user. The notification is provided both visually and through voice guidance using speech synthesis technology. The input is the solution sent from the server, and the output is the visual and voice notification to the user. This allows the user to receive specific instructions and confidently take the next action.

[0343] (Application Example 2)

[0344] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0345] When elderly people use communication devices in commercial facilities and other locations, they often find them difficult to operate due to their unfamiliarity with the technology, which can increase stress and anxiety. This can also impair their user experience and reduce convenience. There is a need to provide systems that allow elderly people to operate communication devices with greater confidence and enjoy a more fulfilling user experience.

[0346] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0347] In this invention, the server includes means for analyzing visual displays installed on the user's communication device, means for identifying the problem the user is facing, means for presenting solutions, and means for recognizing the user's emotional state and adjusting the solutions. This makes it possible for elderly people to use communication devices safely and without stress in commercial facilities and other places.

[0348] A "communication device" is a device used for voice communication, data communication, etc., and is a device that allows users to input and receive information.

[0349] "Analysis means" refers to a device or program that has the function of analyzing visual displays or input data and understanding or identifying their content.

[0350] "Identification means" refers to the technology or process used to identify and clarify the challenges faced by users, using analyzed data.

[0351] A "presentation method" is a system that provides users with appropriate solutions or information based on problems identified through specific means.

[0352] "Emotion recognition means" refers to technologies and methods that analyze a user's facial expressions, voice, and actions to identify and evaluate their psychological state and emotions.

[0353] The system that realizes this invention will enable elderly people to safely obtain information using communication devices in commercial facilities and other locations. The server will function via an application installed on the communication device and will utilize various hardware and software.

[0354] The communication device includes hardware such as cameras, microphones, speakers, and displays. The analysis method uses this hardware to acquire image data and analyze the visual display. Software used includes computer vision technology (e.g., OCR technology) and emotion recognition technology (e.g., machine learning algorithms). Specifically, for image analysis, services such as Amazon Rekognition and Google Cloud Vision can be applied, and for emotion recognition, methods that read emotions from voice and facial expressions are employed.

[0355] The server processes data collected by the analysis means and identifies the user's challenges using the identification means. Based on these identified challenges, the presentation means presents appropriate solutions and information to the user via voice or text display. The emotion recognition means understands the user's psychological state and adjusts the solutions according to the user's emotions. As a result, elderly people can operate communication devices smoothly and without stress when using them in commercial facilities.

[0356] As a concrete example, if a user is lost in the store, the camera on the communication device reads and analyzes the store's map information. If the user appears frustrated, a gentle voice message such as, "Hello! The special sale section is over here. Do you need any guidance?" and a simple map display are provided.

[0357] Examples of input prompts for a generative AI model are as follows:

[0358] "Please create a message that will appropriately guide elderly customers who are lost in the store. It should be especially reassuring for those who are unsure about how to use the machines."

[0359] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0360] Step 1:

[0361] The device uses its camera to acquire visual information about its surroundings. This input data is in image format and is transmitted to the analysis system. The device uses OCR technology to convert the image data into text data and recognize store maps and product labels.

[0362] Step 2:

[0363] The device uses identified text data and collected location information (e.g., GPS or in-store beacons) to determine the user's current location and what they are looking at. This data is sent to a server to identify the user's location and interests.

[0364] Step 3:

[0365] The server analyzes the data sent from the terminal and forms prompt sentences to input into the AI ​​model. These prompt sentences are adjusted based on the user's current location, interests, and past behavioral history. Furthermore, the server evaluates the user's challenges and stress levels based on data from emotion recognition systems.

[0366] Step 4:

[0367] The server requests a solution from the AI ​​model based on the prompt message and generates a guidance message as output. For example, it might create a message such as, "The special sale section is over here. Do you need guidance?"

[0368] Step 5:

[0369] The generated guidance message is sent to the terminal and visually displayed to the user through a display device, as well as provided as voice guidance using speech synthesis technology. This allows the user to obtain guidance information from their current location to their destination, supporting them in traveling without stress.

[0370] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0371] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0372] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0373] [Third Embodiment]

[0374] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0375] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0376] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0377] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0378] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0379] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0380] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0381] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0382] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0383] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0384] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0385] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0386] This invention is an AI assistant system that helps elderly people use smartphones more easily. This system is installed as software on the user's smartphone, which is the user's operating terminal, and analyzes visual information on the screen to identify problems the user is facing.

[0387] First, the device uses its built-in screen capture function to constantly monitor the screen content. The acquired visual information is analyzed as text using OCR (Optical Character Recognition) technology, and images and patterns are recognized through computer vision. This allows the device to classify different elements such as advertisements, alerts, or messages.

[0388] Next, the device monitors the user's touch input and operation logs, identifying which operations the user is having difficulty with based on delays and multiple attempts at specific operations. This information is sent to a server, which uses machine learning algorithms to analyze it and generate solutions needed by the user. These solutions are expressed in concise and clear language to aid the user's understanding.

[0389] The generated solution is displayed as a visual notification on the device or provided to the user as an audio message using speech synthesis technology. For example, if a user cannot find the delete button in an email app, the device will instruct them via audio or on-screen message, "The delete button is in the upper right corner." This embodiment enables the rapid and effective resolution of problems faced by users and provides an environment where even elderly people can easily use smartphones.

[0390] As a concrete example of this system, consider a scenario where an advertisement appears while a user is using a social media app. In this situation, the device detects the advertisement and notifies the user with a message such as, "This is an advertisement. No action is required," supporting the user so they can continue using the app without confusion. This allows users to use digital devices with greater peace of mind.

[0391] The following describes the processing flow.

[0392] Step 1:

[0393] The device periodically captures the user's smartphone screen using its built-in screen capture function. This captured data is used as visual information for subsequent analysis.

[0394] Step 2:

[0395] The device analyzes the visual information of the captured screen using OCR (Optical Character Recognition) technology. This extracts the text information on the screen as text data. Furthermore, it utilizes computer vision to recognize non-text elements, including images and icons.

[0396] Step 3:

[0397] The device sends the acquired text and image data to the server. The server receives the data and applies machine learning algorithms to classify different elements such as advertisements, alerts, and messages. This classification creates a foundation for understanding what information the user is interacting with.

[0398] Step 4:

[0399] The device collects user touch events and interaction logs. If a user repeatedly taps a specific button or stays on the same screen for an extended period, it detects and logs this. This helps determine if the user may be experiencing difficulties.

[0400] Step 5:

[0401] The server identifies the problems the user is facing based on the operation logs and classification data received from the terminal. For example, it may identify cases where the user has difficulty distinguishing between advertisements and content, or where the user does not understand the operating procedure.

[0402] Step 6:

[0403] The server generates a solution to the problem and creates a message to provide to the user in a concise and easy-to-understand format. This message is output as text that also serves as audio guidance, along with visual instructions.

[0404] Step 7:

[0405] The terminal notifies the user of a solution message received from the server. This can be done by displaying a pop-up message on the screen or by providing voice instructions to the user using speech synthesis technology.

[0406] Step 8:

[0407] The user follows instructions from the terminal and attempts to resolve the problem. After resolving the issue, the terminal monitors the user's actions again and sends the results as feedback to the server. This feedback is used to improve the system and enhance its accuracy.

[0408] (Example 1)

[0409] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0410] The present invention aims to provide a technology that effectively solves the operational difficulties that arise when users, including the elderly, use operating terminals due to the large amount of information on the screen and complex operations. Conventional technologies have made it difficult to identify users' operational problems in real time and provide appropriate support immediately. As a result, this has often compromised the convenience and sense of security of users.

[0411] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0412] In this invention, the server includes an analysis means installed on the operating device for analyzing image information displayed on the display surface, a monitoring means for monitoring user input operations and operation history and identifying user difficulties from delays and multiple attempts, and a notification means for notifying the user of the generated solution using visual or speech synthesis technology. This enables the rapid identification of problems in the user's operation and the provision of appropriate solutions.

[0413] An "operating device" is an electronic terminal device that users can directly operate, and it serves as the foundation on which this system is installed.

[0414] "Analysis means" refers to means that have the function of analyzing image information displayed on a display surface and identifying text data and other patterns.

[0415] "Identification means" refers to a means of identifying the problem a user is facing using information obtained through analysis means.

[0416] A "presentation means" is a means that has the function of providing a solution to the user based on the problem identified by the identification means.

[0417] "Monitoring means" refers to methods for monitoring user input operations and operation history in real time to identify delays or errors in operations.

[0418] A "generation method" is a means of sending information collected by a monitoring method to a server and using machine learning to create the optimal solution for the user.

[0419] A "notification method" is a means of communicating the generated solution to the user, and it uses visual displays or speech synthesis technology to notify the information.

[0420] This invention relates to an assistance system installed on electronic terminals operated by users, particularly the elderly. Embodiments of this system are described below.

[0421] A terminal refers to an electronic device with a user interface, such as a smartphone or tablet, where a program is executed. The terminal has a built-in screen capture function, which periodically captures information from the display screen. The captured images are analyzed using OCR (Optical Character Recognition) software and computer vision technology to identify text data and other visual information. This allows, for example, determining how advertisements or notification messages are displayed on the screen.

[0422] The terminal also features transparent monitoring capabilities, collecting user touch input and operation logs. This data provides crucial clues for analyzing where users are experiencing difficulties with operation. For example, if a particular button is frequently pressed incorrectly, the terminal records that operation and requests an appropriate solution from the server.

[0423] The server receives information sent from the terminal and analyzes the problem using machine learning algorithms. It then uses a generative AI model to generate solutions tailored to individual users. These solutions are provided to the user on the terminal via visual notifications or voice messages using speech synthesis technology. For example, clear instructions such as "The delete button is in the upper right corner" may be given.

[0424] For example, if an advertisement appears while a user is using a social media app, the device recognizes the advertisement and notifies the user with a message saying, "This is an advertisement. No action is required." In this way, users can avoid confusion caused by unnecessary actions and continue to use electronic devices with peace of mind.

[0425] As an example of a prompt, the model might be given an instruction such as, "Please suggest solutions to make smartphones easier for the elderly to use," and appropriate support content will be generated accordingly.

[0426] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0427] Step 1:

[0428] The device periodically captures screen content using its built-in screen capture function. The input consists of all screen information displayed on the device's screen. The captured images are then directly sent as output for OCR and computer vision analysis. This process is performed every few seconds, allowing the user to evaluate the applications and notifications they are currently using.

[0429] Step 2:

[0430] The device converts the acquired screenshot into text data using OCR technology and then uses computer vision to identify images and patterns. The input is a screenshot image, and the output is text data and associated pattern information. Based on this information, it identifies elements such as advertisements and icons within apps. As a specific example, it identifies advertising banners displayed on social media apps.

[0431] Step 3:

[0432] The device monitors the user's touch input and operation history, and records the information. Input includes the user's screen touch location and time, and output is a log of that operation history. If repeated errors occur over a certain period, the data is used to identify difficult operations. Specifically, this involves counting the number of times the same icon is tapped repeatedly.

[0433] Step 4:

[0434] The device sends information about the identified user's operational problems to the server. The input is data about the identified operation history, and the output is a request to the server. This is sent to the server for further analysis. Specifically, it sends data on the number of times the user has made a mistake deleting operation in the email application.

[0435] Step 5:

[0436] The server analyzes the received data using machine learning algorithms. The input is operation data sent from the terminal, and the output is a solution based on that data. A generative AI model is used to determine specific assistance. For example, it might generate guidance such as, "The delete button is in the upper right corner of the screen."

[0437] Step 6:

[0438] The device notifies the user of the solution received from the server. The input is the solution data from the server, and the output is the notification to the user. The notification is either displayed visually on the screen or played back using speech synthesis technology. For example, the user is informed via voice, "You don't need to worry about the advertisements."

[0439] (Application Example 1)

[0440] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0441] When elderly people use digital devices, especially when dealing with complex electronic payment systems, errors or inaccurate judgments can lead to problems. These challenges pose a barrier to the safe and secure use of electronic payments. Therefore, there is a need to provide an environment that allows users to make electronic payments more intuitively and securely.

[0442] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0443] In this invention, the server includes analysis means for analyzing visual information installed on the operating device, identification means for identifying a problem faced by the user using the obtained information, notification means for presenting a solution to the user based on the identified problem, and support means for monitoring transaction screens and guiding the user to points where decision-making is required. This enables the user to proceed with complex electronic payment operations with confidence by following the guidance.

[0444] An "operating device" is an electronic device used by the user to perform operations, and in this context, it mainly refers to a smartphone.

[0445] "Visual information" refers to visual data such as text and images displayed on a screen.

[0446] "Analysis means" refers to technical means used to read visual information and understand its content.

[0447] "Information" refers to data and insights obtained through analytical methods.

[0448] An "identification method" is a function that identifies the challenges and problems that the user faces based on the information obtained.

[0449] "Notification method" refers to a method of informing the user of solutions to identified problems.

[0450] "Support measures" refer to functions that monitor transaction-related screens and guide users to points where decision-making is required.

[0451] "Transaction-related screens" refer to interface screens where processes such as electronic payments are processed.

[0452] "Decision-making" refers to the judgments that users need to make when proceeding with transactions or operations.

[0453] The system for realizing this invention is installed on the user's operating device and consists mainly of a program that processes various types of information. The server constantly monitors the visual information on the screen through analysis means running on the operating device and recognizes it as text using OCR (optical character recognition) technology. Furthermore, it uses computer vision to classify images and patterns and identify screens related to transactions.

[0454] The analysis means utilizes this visual information to help the identification means identify the user's problem. Based on the identified problem, the notification means uses speech synthesis technology to notify the user of appropriate actions or instructions from the server. For example, it provides clear guidance to the user regarding the location of the button for completing a transaction and the steps involved in its operation.

[0455] Furthermore, the support system monitors the transaction selection screen when the user is making an electronic payment, providing voice and on-screen instructions at appropriate points where the user needs to make a decision. This allows the user to continue the operation safely.

[0456] As a concrete example, consider a scenario where an elderly user is making an online purchase. In this case, a voice prompt such as, "Tap this button to complete the payment," would be provided, allowing the user to complete the purchase without confusion.

[0457] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0458] Step 1:

[0459] The device periodically acquires visual information from the screen using its screen capture function. The input is the visual information displayed on the device's screen, and the output is the storage of this information as image data. This acquired data is then converted into text data using OCR technology.

[0460] Step 2:

[0461] The server analyzes the acquired text data and uses identification methods to identify potential problems the user may be facing. The input is text data converted by OCR, and the output is a clarification of the issues the user needs to address. It analyzes words and phrases within the text to determine which information is important.

[0462] Step 3:

[0463] The device presents appropriate solutions to identified problems. It uses notification methods and speech synthesis technology to present solutions to the user. The input is the clearly defined problem the user faces, and the output is voice guidance or on-screen instructions. For example, it provides specific voice instructions such as, "Tap this button to complete the payment."

[0464] Step 4:

[0465] The support system monitors transaction-related screens and assists users in performing transactions. Input is the state of the transaction screen, and output is user guidance regarding necessary actions. It recognizes specific elements of the transaction screen and generates messages prompting the user to take the next action.

[0466] Step 5:

[0467] The user makes decisions by following the instructions provided by the device. The input is the instructional messages from the device, and the output is the user's choice of action. Based on the information provided, the user can proceed with the necessary actions without hesitation.

[0468] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0469] This invention is an AI assistant system that helps elderly people use smartphones more effectively, and in particular, incorporates a function to recognize the user's emotions. The system is installed on the user's terminal, analyzes various data inputs in real time, identifies the problems they face, and then presents solutions that are tailored to their emotions.

[0470] First, the device continuously collects visual information on the screen using its screen capture function. This data is analyzed using OCR technology, recognizing image elements through computer vision. This allows the information on the screen to be analyzed and classified into elements such as advertisements, alerts, and general messages.

[0471] Next, the device senses the user's voice and facial expressions and analyzes those emotions using its built-in emotion engine. By combining data such as voice tone, changes in facial expressions, and even operation speed, it infers the user's emotional state, such as whether they are feeling stressed or comfortable using the device.

[0472] This information is sent to a server, which uses machine learning algorithms to analyze the user's current situation and emotional state. Based on these results, the server identifies the specific problem the user is experiencing and generates an appropriate solution. The solution is tailored to the user's emotions and is designed to provide reassurance.

[0473] The generated solutions are sent to the device and notified to the user. These notifications include both on-screen visuals and voice guidance utilizing speech synthesis technology. For example, if a user becomes confused and frustrated while using the app, the system provides an encouraging message such as, "Try this, you'll get the hang of it." In this way, the system supports older adults in using technology with confidence through user-emotion-based feedback.

[0474] The following describes the processing flow.

[0475] Step 1:

[0476] The device uses a screen capture function to periodically capture the user's smartphone screen and collect data as visual information.

[0477] Step 2:

[0478] The device uses OCR technology to extract text from captured visual information, and further recognizes and classifies images and icons using computer vision.

[0479] Step 3:

[0480] The device uses a camera and microphone to detect the user's voice and facial expressions, and an emotion engine analyzes their emotional state. This analysis determines the user's stress levels, happiness, and other emotional states.

[0481] Step 4:

[0482] The terminal sends these analysis results to the server, which uses the received data to identify the problems the user is facing using machine learning algorithms.

[0483] Step 5:

[0484] The server generates customized solutions to resolve problems based on the user's emotional state. These solutions are designed as messages that are considerate and reassuring to the user.

[0485] Step 6:

[0486] The terminal receives the solution sent from the server and notifies the user. This notification is provided both as a visual message on the UI and as voice guidance using speech synthesis technology.

[0487] Step 7:

[0488] The user attempts to resolve the problem by following the instructions provided by the terminal. The terminal collects the results of these operations as feedback and sends it to the server. This feedback is used to improve the system.

[0489] (Example 2)

[0490] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0491] In modern society, elderly people face difficulties in effectively using smartphones. The complexity of operation and information overload can cause stress and lead to overlooking important information. Therefore, there is a need for systems that provide operational support and emotionally sensitive responses to enable elderly people to use smartphones with peace of mind.

[0492] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0493] In this invention, the server includes a collection means for continuously collecting the user's visual information and analyzing textual information, a classification means for classifying the visual information, and an emotion recognition means for analyzing the emotional state. This makes it possible to identify the problems the user is facing and provide solutions that give a sense of security in line with their emotions.

[0494] "Collection means" refers to a function on the user's terminal that continuously collects visual information displayed on the screen and analyzes the textual information using optical character recognition technology.

[0495] A "classification method" is a function that uses computer vision technology to analyze visual information acquired by collection methods and identify and classify elements such as advertisements, alerts, and general messages.

[0496] "Emotion recognition means" refers to a function that senses the user's voice and facial expressions, and infers the user's emotional state by analyzing voice tone, facial expression changes, and the speed of a series of movements.

[0497] "Identification means" refers to a function that uses data obtained from classification means and emotion recognition means to identify specific problems and challenges that the user is facing.

[0498] The "generation method" refers to a function that applies machine learning algorithms to data sent to the server to generate appropriate solutions. The solutions are adjusted according to the user's emotional state.

[0499] "Presentation means" refers to a function that provides users with solutions created by the generation means using visual display or speech synthesis technology.

[0500] This invention is a support system for elderly people to use smartphones more effectively, providing an AI assistant combined with emotion recognition capabilities. This system is installed on the user's terminal and utilizes hardware such as screen capture functionality, microphones, and cameras built into the terminal, as well as software such as OCR technology and computer vision technology.

[0501] The device collects information on the screen in real time using its screen capture function and analyzes that information using OCR technology. The data extracted through this analysis is then classified into elements such as advertisements, alerts, and general messages on the screen using computer vision technology. In addition, the device senses the user's voice and facial expressions using a microphone and camera, and analyzes the voice tone and facial expressions to infer the user's emotional state.

[0502] These analysis results are sent to a server, which uses machine learning algorithms to analyze the user's current situation and emotional state. Based on this analysis, the server generates an appropriate solution and adjusts it according to the user's emotions. For example, if the user is feeling anxious, the server will generate a message using reassuring language. The generated solution is sent to the terminal and notified to the user. The terminal can not only display the solution visually on the screen but can also provide voice guidance using speech synthesis technology.

[0503] Specifically, when a user gets lost in the app's settings, they will be given specific voice guidance such as, "Tap the gear icon here, and then select 'Network Settings'." In this way, users can instantly receive feedback that matches their emotional state at that moment.

[0504] An example of a prompt might be, "If a user is struggling to understand a new feature in the app, how can we guide them to a solution while also providing reassurance?" This prompt allows the generative AI model to derive and present individual solutions tailored to the user's situation.

[0505] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0506] Step 1:

[0507] The device collects visual information on the screen in real time using its screen capture function. This collected data is then analyzed using OCR technology to extract text data. The input is the visual information on the screen, and the output is the analyzed text data. In this process, the device prepares the basic data necessary to identify elements such as advertisements, alerts, and general messages.

[0508] Step 2:

[0509] The device uses computer vision technology to analyze image elements on the screen and combine them with collected text data to classify advertisements, alerts, general messages, etc. The input is text data and image data extracted by OCR, and the output is classified visual information. This classification process clarifies the type of information the user is facing on the screen.

[0510] Step 3:

[0511] The device uses a microphone and camera to sense the user's voice and facial expressions. This allows for the collection of voice tone and facial expression data. The input is the user's real-time voice and facial expressions captured by the microphone and camera, and the output is data for emotion estimation. The device uses this data to infer the user's emotional state.

[0512] Step 4:

[0513] The device sends collected visual and emotional data to the server. The server analyzes this data using machine learning algorithms. The input is the classified visual and emotional data sent from the device, and the output is the analyzed user situation and emotional state. Based on this analysis, the server identifies the problems the user is facing.

[0514] Step 5:

[0515] Based on the identified problem, the server uses a generative AI model to generate appropriate solutions that are sensitive to the user's emotions. The input is the server's analysis of the user's situation and problem, and the output is a customized solution. The generated solution is adjusted according to the user's emotional state.

[0516] Step 6:

[0517] The server sends the generated solution to the terminal. The terminal receives this solution and notifies the user. The notification is provided both visually and through voice guidance using speech synthesis technology. The input is the solution sent from the server, and the output is the visual and voice notification to the user. This allows the user to receive specific instructions and confidently take the next action.

[0518] (Application Example 2)

[0519] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0520] When elderly people use communication devices in commercial facilities and other locations, they often find them difficult to operate due to their unfamiliarity with the technology, which can increase stress and anxiety. This can also impair their user experience and reduce convenience. There is a need to provide systems that allow elderly people to operate communication devices with greater confidence and enjoy a more fulfilling user experience.

[0521] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0522] In this invention, the server includes means for analyzing visual displays installed on the user's communication device, means for identifying the problem the user is facing, means for presenting solutions, and means for recognizing the user's emotional state and adjusting the solutions. This makes it possible for elderly people to use communication devices safely and without stress in commercial facilities and other places.

[0523] A "communication device" is a device used for voice communication, data communication, etc., and is a device that allows users to input and receive information.

[0524] "Analysis means" refers to a device or program that has the function of analyzing visual displays or input data and understanding or identifying their content.

[0525] "Identification means" refers to the technology or process used to identify and clarify the challenges faced by users, using analyzed data.

[0526] A "presentation method" is a system that provides users with appropriate solutions or information based on problems identified through specific means.

[0527] "Emotion recognition means" refers to technologies and methods that analyze a user's facial expressions, voice, and actions to identify and evaluate their psychological state and emotions.

[0528] The system that realizes this invention will enable elderly people to safely obtain information using communication devices in commercial facilities and other locations. The server will function via an application installed on the communication device and will utilize various hardware and software.

[0529] The communication device includes hardware such as cameras, microphones, speakers, and displays. The analysis method uses this hardware to acquire image data and analyze the visual display. Software used includes computer vision technology (e.g., OCR technology) and emotion recognition technology (e.g., machine learning algorithms). Specifically, for image analysis, services such as Amazon Rekognition and Google Cloud Vision can be applied, and for emotion recognition, methods that read emotions from voice and facial expressions are employed.

[0530] The server processes data collected by the analysis means and identifies the user's challenges using the identification means. Based on these identified challenges, the presentation means presents appropriate solutions and information to the user via voice or text display. The emotion recognition means understands the user's psychological state and adjusts the solutions according to the user's emotions. As a result, elderly people can operate communication devices smoothly and without stress when using them in commercial facilities.

[0531] As a concrete example, if a user is lost in the store, the camera on the communication device reads and analyzes the store's map information. If the user appears frustrated, a gentle voice message such as, "Hello! The special sale section is over here. Do you need any guidance?" and a simple map display are provided.

[0532] Examples of input prompts for a generative AI model are as follows:

[0533] "Please create a message that will appropriately guide elderly customers who are lost in the store. It should be especially reassuring for those who are unsure about how to use the machines."

[0534] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0535] Step 1:

[0536] The device uses its camera to acquire visual information about its surroundings. This input data is in image format and is transmitted to the analysis system. The device uses OCR technology to convert the image data into text data and recognize store maps and product labels.

[0537] Step 2:

[0538] The device uses identified text data and collected location information (e.g., GPS or in-store beacons) to determine the user's current location and what they are looking at. This data is sent to a server to identify the user's location and interests.

[0539] Step 3:

[0540] The server analyzes the data sent from the terminal and forms prompt sentences to input into the AI ​​model. These prompt sentences are adjusted based on the user's current location, interests, and past behavioral history. Furthermore, the server evaluates the user's challenges and stress levels based on data from emotion recognition systems.

[0541] Step 4:

[0542] The server requests a solution from the AI ​​model based on the prompt message and generates a guidance message as output. For example, it might create a message such as, "The special sale section is over here. Do you need guidance?"

[0543] Step 5:

[0544] The generated guidance message is sent to the terminal and visually displayed to the user through a display device, as well as provided as voice guidance using speech synthesis technology. This allows the user to obtain guidance information from their current location to their destination, supporting them in traveling without stress.

[0545] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0546] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0547] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0548] [Fourth Embodiment]

[0549] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0550] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0551] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0552] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0553] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0554] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0555] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0556] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0557] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0558] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0559] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0560] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0561] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0562] This invention is an AI assistant system that helps elderly people use smartphones more easily. This system is installed as software on the user's smartphone, which is the user's operating terminal, and analyzes visual information on the screen to identify problems the user is facing.

[0563] First, the device uses its built-in screen capture function to constantly monitor the screen content. The acquired visual information is analyzed as text using OCR (Optical Character Recognition) technology, and images and patterns are recognized through computer vision. This allows the device to classify different elements such as advertisements, alerts, or messages.

[0564] Next, the device monitors the user's touch input and operation logs, identifying which operations the user is having difficulty with based on delays and multiple attempts at specific operations. This information is sent to a server, which uses machine learning algorithms to analyze it and generate solutions needed by the user. These solutions are expressed in concise and clear language to aid the user's understanding.

[0565] The generated solution is displayed as a visual notification on the device or provided to the user as an audio message using speech synthesis technology. For example, if a user cannot find the delete button in an email app, the device will instruct them via audio or on-screen message, "The delete button is in the upper right corner." This embodiment enables the rapid and effective resolution of problems faced by users and provides an environment where even elderly people can easily use smartphones.

[0566] As a concrete example of this system, consider a scenario where an advertisement appears while a user is using a social media app. In this situation, the device detects the advertisement and notifies the user with a message such as, "This is an advertisement. No action is required," supporting the user so they can continue using the app without confusion. This allows users to use digital devices with greater peace of mind.

[0567] The following describes the processing flow.

[0568] Step 1:

[0569] The device periodically captures the user's smartphone screen using its built-in screen capture function. This captured data is used as visual information for subsequent analysis.

[0570] Step 2:

[0571] The device analyzes the visual information of the captured screen using OCR (Optical Character Recognition) technology. This extracts the text information on the screen as text data. Furthermore, it utilizes computer vision to recognize non-text elements, including images and icons.

[0572] Step 3:

[0573] The device sends the acquired text and image data to the server. The server receives the data and applies machine learning algorithms to classify different elements such as advertisements, alerts, and messages. This classification creates a foundation for understanding what information the user is interacting with.

[0574] Step 4:

[0575] The device collects user touch events and interaction logs. If a user repeatedly taps a specific button or stays on the same screen for an extended period, it detects and logs this. This helps determine if the user may be experiencing difficulties.

[0576] Step 5:

[0577] The server identifies the problems the user is facing based on the operation logs and classification data received from the terminal. For example, it may identify cases where the user has difficulty distinguishing between advertisements and content, or where the user does not understand the operating procedure.

[0578] Step 6:

[0579] The server generates a solution to the problem and creates a message to provide to the user in a concise and easy-to-understand format. This message is output as text that also serves as audio guidance, along with visual instructions.

[0580] Step 7:

[0581] The terminal notifies the user of a solution message received from the server. This can be done by displaying a pop-up message on the screen or by providing voice instructions to the user using speech synthesis technology.

[0582] Step 8:

[0583] The user follows instructions from the terminal and attempts to resolve the problem. After resolving the issue, the terminal monitors the user's actions again and sends the results as feedback to the server. This feedback is used to improve the system and enhance its accuracy.

[0584] (Example 1)

[0585] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0586] The present invention aims to provide a technology that effectively solves the operational difficulties that arise when users, including the elderly, use operating terminals due to the large amount of information on the screen and complex operations. Conventional technologies have made it difficult to identify users' operational problems in real time and provide appropriate support immediately. As a result, this has often compromised the convenience and sense of security of users.

[0587] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0588] In this invention, the server includes an analysis means installed on the operating device for analyzing image information displayed on the display surface, a monitoring means for monitoring user input operations and operation history and identifying user difficulties from delays and multiple attempts, and a notification means for notifying the user of the generated solution using visual or speech synthesis technology. This enables the rapid identification of problems in the user's operation and the provision of appropriate solutions.

[0589] An "operating device" is an electronic terminal device that users can directly operate, and it serves as the foundation on which this system is installed.

[0590] "Analysis means" refers to means that have the function of analyzing image information displayed on a display surface and identifying text data and other patterns.

[0591] "Identification means" refers to a means of identifying the problem a user is facing using information obtained through analysis means.

[0592] A "presentation means" is a means that has the function of providing a solution to the user based on the problem identified by the identification means.

[0593] "Monitoring means" refers to methods for monitoring user input operations and operation history in real time to identify delays or errors in operations.

[0594] A "generation method" is a means of sending information collected by a monitoring method to a server and using machine learning to create the optimal solution for the user.

[0595] A "notification method" is a means of communicating the generated solution to the user, and it uses visual displays or speech synthesis technology to notify the information.

[0596] This invention relates to an assistance system installed on electronic terminals operated by users, particularly the elderly. Embodiments of this system are described below.

[0597] A terminal refers to an electronic device with a user interface, such as a smartphone or tablet, where a program is executed. The terminal has a built-in screen capture function, which periodically captures information from the display screen. The captured images are analyzed using OCR (Optical Character Recognition) software and computer vision technology to identify text data and other visual information. This allows, for example, determining how advertisements or notification messages are displayed on the screen.

[0598] The terminal also features transparent monitoring capabilities, collecting user touch input and operation logs. This data provides crucial clues for analyzing where users are experiencing difficulties with operation. For example, if a particular button is frequently pressed incorrectly, the terminal records that operation and requests an appropriate solution from the server.

[0599] The server receives information sent from the terminal and analyzes the problem using machine learning algorithms. It then uses a generative AI model to generate solutions tailored to individual users. These solutions are provided to the user on the terminal via visual notifications or voice messages using speech synthesis technology. For example, clear instructions such as "The delete button is in the upper right corner" may be given.

[0600] For example, if an advertisement appears while a user is using a social media app, the device recognizes the advertisement and notifies the user with a message saying, "This is an advertisement. No action is required." In this way, users can avoid confusion caused by unnecessary actions and continue to use electronic devices with peace of mind.

[0601] As an example of a prompt, the model might be given an instruction such as, "Please suggest solutions to make smartphones easier for the elderly to use," and appropriate support content will be generated accordingly.

[0602] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0603] Step 1:

[0604] The device periodically captures screen content using its built-in screen capture function. The input consists of all screen information displayed on the device's screen. The captured images are then directly sent as output for OCR and computer vision analysis. This process is performed every few seconds, allowing the user to evaluate the applications and notifications they are currently using.

[0605] Step 2:

[0606] The device converts the acquired screenshot into text data using OCR technology and then uses computer vision to identify images and patterns. The input is a screenshot image, and the output is text data and associated pattern information. Based on this information, it identifies elements such as advertisements and icons within apps. As a specific example, it identifies advertising banners displayed on social media apps.

[0607] Step 3:

[0608] The device monitors the user's touch input and operation history, and records the information. Input includes the user's screen touch location and time, and output is a log of that operation history. If repeated errors occur over a certain period, the data is used to identify difficult operations. Specifically, this involves counting the number of times the same icon is tapped repeatedly.

[0609] Step 4:

[0610] The device sends information about the identified user's operational problems to the server. The input is data about the identified operation history, and the output is a request to the server. This is sent to the server for further analysis. Specifically, it sends data on the number of times the user has made a mistake deleting operation in the email application.

[0611] Step 5:

[0612] The server analyzes the received data using machine learning algorithms. The input is operation data sent from the terminal, and the output is a solution based on that data. A generative AI model is used to determine specific assistance. For example, it might generate guidance such as, "The delete button is in the upper right corner of the screen."

[0613] Step 6:

[0614] The device notifies the user of the solution received from the server. The input is the solution data from the server, and the output is the notification to the user. The notification is either displayed visually on the screen or played back using speech synthesis technology. For example, the user is informed via voice, "You don't need to worry about the advertisements."

[0615] (Application Example 1)

[0616] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0617] When elderly people use digital devices, especially when dealing with complex electronic payment systems, errors or inaccurate judgments can lead to problems. These challenges pose a barrier to the safe and secure use of electronic payments. Therefore, there is a need to provide an environment that allows users to make electronic payments more intuitively and securely.

[0618] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0619] In this invention, the server includes analysis means for analyzing visual information installed on the operating device, identification means for identifying a problem faced by the user using the obtained information, notification means for presenting a solution to the user based on the identified problem, and support means for monitoring transaction screens and guiding the user to points where decision-making is required. This enables the user to proceed with complex electronic payment operations with confidence by following the guidance.

[0620] An "operating device" is an electronic device used by the user to perform operations, and in this context, it mainly refers to a smartphone.

[0621] "Visual information" refers to visual data such as text and images displayed on a screen.

[0622] "Analysis means" refers to technical means used to read visual information and understand its content.

[0623] "Information" refers to data and insights obtained through analytical methods.

[0624] An "identification method" is a function that identifies the challenges and problems that the user faces based on the information obtained.

[0625] "Notification method" refers to a method of informing the user of solutions to identified problems.

[0626] "Support measures" refer to functions that monitor transaction-related screens and guide users to points where decision-making is required.

[0627] "Transaction-related screens" refer to interface screens where processes such as electronic payments are processed.

[0628] "Decision-making" refers to the judgments that users need to make when proceeding with transactions or operations.

[0629] The system for realizing this invention is installed on the user's operating device and consists mainly of a program that processes various types of information. The server constantly monitors the visual information on the screen through analysis means running on the operating device and recognizes it as text using OCR (optical character recognition) technology. Furthermore, it uses computer vision to classify images and patterns and identify screens related to transactions.

[0630] The analysis means utilizes this visual information to help the identification means identify the user's problem. Based on the identified problem, the notification means uses speech synthesis technology to notify the user of appropriate actions or instructions from the server. For example, it provides clear guidance to the user regarding the location of the button for completing a transaction and the steps involved in its operation.

[0631] Furthermore, the support system monitors the transaction selection screen when the user is making an electronic payment, providing voice and on-screen instructions at appropriate points where the user needs to make a decision. This allows the user to continue the operation safely.

[0632] As a concrete example, consider a scenario where an elderly user is making an online purchase. In this case, a voice prompt such as, "Tap this button to complete the payment," would be provided, allowing the user to complete the purchase without confusion.

[0633] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0634] Step 1:

[0635] The device periodically acquires visual information from the screen using its screen capture function. The input is the visual information displayed on the device's screen, and the output is the storage of this information as image data. This acquired data is then converted into text data using OCR technology.

[0636] Step 2:

[0637] The server analyzes the acquired text data and uses identification methods to identify potential problems the user may be facing. The input is text data converted by OCR, and the output is a clarification of the issues the user needs to address. It analyzes words and phrases within the text to determine which information is important.

[0638] Step 3:

[0639] The device presents appropriate solutions to identified problems. It uses notification methods and speech synthesis technology to present solutions to the user. The input is the clearly defined problem the user faces, and the output is voice guidance or on-screen instructions. For example, it provides specific voice instructions such as, "Tap this button to complete the payment."

[0640] Step 4:

[0641] The support system monitors transaction-related screens and assists users in performing transactions. Input is the state of the transaction screen, and output is user guidance regarding necessary actions. It recognizes specific elements of the transaction screen and generates messages prompting the user to take the next action.

[0642] Step 5:

[0643] The user makes decisions by following the instructions provided by the device. The input is the instructional messages from the device, and the output is the user's choice of action. Based on the information provided, the user can proceed with the necessary actions without hesitation.

[0644] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0645] This invention is an AI assistant system that helps elderly people use smartphones more effectively, and in particular, incorporates a function to recognize the user's emotions. The system is installed on the user's terminal, analyzes various data inputs in real time, identifies the problems they face, and then presents solutions that are tailored to their emotions.

[0646] First, the device continuously collects visual information on the screen using its screen capture function. This data is analyzed using OCR technology, recognizing image elements through computer vision. This allows the information on the screen to be analyzed and classified into elements such as advertisements, alerts, and general messages.

[0647] Next, the device senses the user's voice and facial expressions and analyzes those emotions using its built-in emotion engine. By combining data such as voice tone, changes in facial expressions, and even operation speed, it infers the user's emotional state, such as whether they are feeling stressed or comfortable using the device.

[0648] This information is sent to a server, which uses machine learning algorithms to analyze the user's current situation and emotional state. Based on these results, the server identifies the specific problem the user is experiencing and generates an appropriate solution. The solution is tailored to the user's emotions and is designed to provide reassurance.

[0649] The generated solutions are sent to the device and notified to the user. These notifications include both on-screen visuals and voice guidance utilizing speech synthesis technology. For example, if a user becomes confused and frustrated while using the app, the system provides an encouraging message such as, "Try this, you'll get the hang of it." In this way, the system supports older adults in using technology with confidence through user-emotion-based feedback.

[0650] The following describes the processing flow.

[0651] Step 1:

[0652] The device uses a screen capture function to periodically capture the user's smartphone screen and collect data as visual information.

[0653] Step 2:

[0654] The device uses OCR technology to extract text from captured visual information, and further recognizes and classifies images and icons using computer vision.

[0655] Step 3:

[0656] The device uses a camera and microphone to detect the user's voice and facial expressions, and an emotion engine analyzes their emotional state. This analysis determines the user's stress levels, happiness, and other emotional states.

[0657] Step 4:

[0658] The terminal sends these analysis results to the server, which uses the received data to identify the problems the user is facing using machine learning algorithms.

[0659] Step 5:

[0660] The server generates customized solutions to resolve problems based on the user's emotional state. These solutions are designed as messages that are considerate and reassuring to the user.

[0661] Step 6:

[0662] The terminal receives the solution sent from the server and notifies the user. This notification is provided both as a visual message on the UI and as voice guidance using speech synthesis technology.

[0663] Step 7:

[0664] The user attempts to resolve the problem by following the instructions provided by the terminal. The terminal collects the results of these operations as feedback and sends it to the server. This feedback is used to improve the system.

[0665] (Example 2)

[0666] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0667] In modern society, elderly people face difficulties in effectively using smartphones. The complexity of operation and information overload can cause stress and lead to overlooking important information. Therefore, there is a need for systems that provide operational support and emotionally sensitive responses to enable elderly people to use smartphones with peace of mind.

[0668] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0669] In this invention, the server includes a collection means for continuously collecting the user's visual information and analyzing textual information, a classification means for classifying the visual information, and an emotion recognition means for analyzing the emotional state. This makes it possible to identify the problems the user is facing and provide solutions that give a sense of security in line with their emotions.

[0670] "Collection means" refers to a function on the user's terminal that continuously collects visual information displayed on the screen and analyzes the textual information using optical character recognition technology.

[0671] A "classification method" is a function that uses computer vision technology to analyze visual information acquired by collection methods and identify and classify elements such as advertisements, alerts, and general messages.

[0672] "Emotion recognition means" refers to a function that senses the user's voice and facial expressions, and infers the user's emotional state by analyzing voice tone, facial expression changes, and the speed of a series of movements.

[0673] "Identification means" refers to a function that uses data obtained from classification means and emotion recognition means to identify specific problems and challenges that the user is facing.

[0674] The "generation method" refers to a function that applies machine learning algorithms to data sent to the server to generate appropriate solutions. The solutions are adjusted according to the user's emotional state.

[0675] "Presentation means" refers to a function that provides users with solutions created by the generation means using visual display or speech synthesis technology.

[0676] This invention is a support system for elderly people to use smartphones more effectively, providing an AI assistant combined with emotion recognition capabilities. This system is installed on the user's terminal and utilizes hardware such as screen capture functionality, microphones, and cameras built into the terminal, as well as software such as OCR technology and computer vision technology.

[0677] The device collects information on the screen in real time using its screen capture function and analyzes that information using OCR technology. The data extracted through this analysis is then classified into elements such as advertisements, alerts, and general messages on the screen using computer vision technology. In addition, the device senses the user's voice and facial expressions using a microphone and camera, and analyzes the voice tone and facial expressions to infer the user's emotional state.

[0678] These analysis results are sent to a server, which uses machine learning algorithms to analyze the user's current situation and emotional state. Based on this analysis, the server generates an appropriate solution and adjusts it according to the user's emotions. For example, if the user is feeling anxious, the server will generate a message using reassuring language. The generated solution is sent to the terminal and notified to the user. The terminal can not only display the solution visually on the screen but can also provide voice guidance using speech synthesis technology.

[0679] Specifically, when a user gets lost in the app's settings, they will be given specific voice guidance such as, "Tap the gear icon here, and then select 'Network Settings'." In this way, users can instantly receive feedback that matches their emotional state at that moment.

[0680] An example of a prompt might be, "If a user is struggling to understand a new feature in the app, how can we guide them to a solution while also providing reassurance?" This prompt allows the generative AI model to derive and present individual solutions tailored to the user's situation.

[0681] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0682] Step 1:

[0683] The device collects visual information on the screen in real time using its screen capture function. This collected data is then analyzed using OCR technology to extract text data. The input is the visual information on the screen, and the output is the analyzed text data. In this process, the device prepares the basic data necessary to identify elements such as advertisements, alerts, and general messages.

[0684] Step 2:

[0685] The device uses computer vision technology to analyze image elements on the screen and combine them with collected text data to classify advertisements, alerts, general messages, etc. The input is text data and image data extracted by OCR, and the output is classified visual information. This classification process clarifies the type of information the user is facing on the screen.

[0686] Step 3:

[0687] The device uses a microphone and camera to sense the user's voice and facial expressions. This allows for the collection of voice tone and facial expression data. The input is the user's real-time voice and facial expressions captured by the microphone and camera, and the output is data for emotion estimation. The device uses this data to infer the user's emotional state.

[0688] Step 4:

[0689] The device sends collected visual and emotional data to the server. The server analyzes this data using machine learning algorithms. The input is the classified visual and emotional data sent from the device, and the output is the analyzed user situation and emotional state. Based on this analysis, the server identifies the problems the user is facing.

[0690] Step 5:

[0691] Based on the identified problem, the server uses a generative AI model to generate appropriate solutions that are sensitive to the user's emotions. The input is the server's analysis of the user's situation and problem, and the output is a customized solution. The generated solution is adjusted according to the user's emotional state.

[0692] Step 6:

[0693] The server sends the generated solution to the terminal. The terminal receives this solution and notifies the user. The notification is provided both visually and through voice guidance using speech synthesis technology. The input is the solution sent from the server, and the output is the visual and voice notification to the user. This allows the user to receive specific instructions and confidently take the next action.

[0694] (Application Example 2)

[0695] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0696] When elderly people use communication devices in commercial facilities and other locations, they often find them difficult to operate due to their unfamiliarity with the technology, which can increase stress and anxiety. This can also impair their user experience and reduce convenience. There is a need to provide systems that allow elderly people to operate communication devices with greater confidence and enjoy a more fulfilling user experience.

[0697] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0698] In this invention, the server includes means for analyzing visual displays installed on the user's communication device, means for identifying the problem the user is facing, means for presenting solutions, and means for recognizing the user's emotional state and adjusting the solutions. This makes it possible for elderly people to use communication devices safely and without stress in commercial facilities and other places.

[0699] A "communication device" is a device used for voice communication, data communication, etc., and is a device that allows users to input and receive information.

[0700] "Analysis means" refers to a device or program that has the function of analyzing visual displays or input data and understanding or identifying their content.

[0701] "Identification means" refers to the technology or process used to identify and clarify the challenges faced by users, using analyzed data.

[0702] A "presentation method" is a system that provides users with appropriate solutions or information based on problems identified through specific means.

[0703] "Emotion recognition means" refers to technologies and methods that analyze a user's facial expressions, voice, and actions to identify and evaluate their psychological state and emotions.

[0704] The system that realizes this invention will enable elderly people to safely obtain information using communication devices in commercial facilities and other locations. The server will function via an application installed on the communication device and will utilize various hardware and software.

[0705] The communication device includes hardware such as cameras, microphones, speakers, and displays. The analysis method uses this hardware to acquire image data and analyze the visual display. Software used includes computer vision technology (e.g., OCR technology) and emotion recognition technology (e.g., machine learning algorithms). Specifically, for image analysis, services such as Amazon Rekognition and Google Cloud Vision can be applied, and for emotion recognition, methods that read emotions from voice and facial expressions are employed.

[0706] The server processes data collected by the analysis means and identifies the user's challenges using the identification means. Based on these identified challenges, the presentation means presents appropriate solutions and information to the user via voice or text display. The emotion recognition means understands the user's psychological state and adjusts the solutions according to the user's emotions. As a result, elderly people can operate communication devices smoothly and without stress when using them in commercial facilities.

[0707] As a concrete example, if a user is lost in the store, the camera on the communication device reads and analyzes the store's map information. If the user appears frustrated, a gentle voice message such as, "Hello! The special sale section is over here. Do you need any guidance?" and a simple map display are provided.

[0708] Examples of input prompts for a generative AI model are as follows:

[0709] "Please create a message that will appropriately guide elderly customers who are lost in the store. It should be especially reassuring for those who are unsure about how to use the machines."

[0710] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0711] Step 1:

[0712] The device uses its camera to acquire visual information about its surroundings. This input data is in image format and is transmitted to the analysis system. The device uses OCR technology to convert the image data into text data and recognize store maps and product labels.

[0713] Step 2:

[0714] The device uses identified text data and collected location information (e.g., GPS or in-store beacons) to determine the user's current location and what they are looking at. This data is sent to a server to identify the user's location and interests.

[0715] Step 3:

[0716] The server analyzes the data sent from the terminal and forms prompt sentences to input into the AI ​​model. These prompt sentences are adjusted based on the user's current location, interests, and past behavioral history. Furthermore, the server evaluates the user's challenges and stress levels based on data from emotion recognition systems.

[0717] Step 4:

[0718] The server requests a solution from the AI ​​model based on the prompt message and generates a guidance message as output. For example, it might create a message such as, "The special sale section is over here. Do you need guidance?"

[0719] Step 5:

[0720] The generated guidance message is sent to the terminal and visually displayed to the user through a display device, as well as provided as voice guidance using speech synthesis technology. This allows the user to obtain guidance information from their current location to their destination, supporting them in traveling without stress.

[0721] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0722] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0723] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0724] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0725] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0726] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0727] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0728] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0729] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0730] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0731] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0732] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0733] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0734] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0735] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0736] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0737] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0738] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0739] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0740] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0741] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0742] The following is further disclosed regarding the embodiments described above.

[0743] (Claim 1)

[0744] An analysis means installed on the user's terminal and used to analyze visual information displayed on the screen,

[0745] A means for identifying problems faced by users using data obtained by the said analysis means,

[0746] A means of presenting solutions to the user based on the problem identified by the specified means,

[0747] A system that includes this.

[0748] (Claim 2)

[0749] Furthermore, it includes monitoring means to monitor user actions and continuously collect data.

[0750] The system according to claim 1, wherein the data acquired by the monitoring means is used by the identifying means.

[0751] (Claim 3)

[0752] The system according to claim 1, wherein the presentation means provides the solution as voice using speech synthesis technology.

[0753] "Example 1"

[0754] (Claim 1)

[0755] An analysis means installed in the operating device and used to analyze image information displayed on the display surface,

[0756] An identification means that identifies the problem the user is facing using the information obtained by the analysis means,

[0757] A presentation means that presents solutions to the user based on the problem identified by the identification means,

[0758] A monitoring system that monitors user input operations and operation history, and identifies user difficulties from delays and multiple attempts,

[0759] A generation means that transmits information collected by the monitoring means to a server, and the server generates a solution using machine learning,

[0760] A notification means for informing the user of the generated solution using visual or speech synthesis technology,

[0761] A system that includes this.

[0762] (Claim 2)

[0763] The system according to claim 1, which continuously collects operational information and is used by the analysis means and the identification means.

[0764] (Claim 3)

[0765] The system according to claim 1, wherein the notification means provides a solution in voice using speech synthesis technology.

[0766] "Application Example 1"

[0767] (Claim 1)

[0768] An analysis means installed on the user's operating device to analyze visual information,

[0769] An identification means that identifies the challenges faced by the user using the information obtained by the analysis means,

[0770] A notification means that presents a solution to the user based on the problem identified by the identification means,

[0771] A support system that monitors transaction-related screens and guides users to points where decision-making is required,

[0772] A system that includes this.

[0773] (Claim 2)

[0774] Furthermore, it is equipped with observation means to monitor user operations and continuously collect information,

[0775] The system according to claim 1, wherein the identification means uses the information obtained by the observation means.

[0776] (Claim 3)

[0777] The system according to claim 1, wherein the notification means provides the solution as voice using voice generation technology.

[0778] "Example 2 of combining an emotion engine"

[0779] (Claim 1)

[0780] A collection means installed on the user's operating terminal that continuously collects visual information displayed on the screen and analyzes textual information,

[0781] A classification means for classifying the visual information acquired by the said collection means using computer vision technology,

[0782] An emotion recognition means that senses the user's voice and facial expressions and analyzes their emotional state,

[0783] An identification means that identifies the problem the user is facing using the data acquired by the classification means and the emotion recognition means,

[0784] A generation means that analyzes data sent to a server using a machine learning algorithm and generates solutions tailored to the user,

[0785] A presentation means that presents the solution generated by the said generation means and provides it using visual display and speech synthesis technology,

[0786] A system that includes this.

[0787] (Claim 2)

[0788] Furthermore, it includes monitoring mechanisms to monitor the user's voice tone and operation speed, and to infer their emotional state.

[0789] The system according to claim 1, wherein the data acquired by the monitoring means is used by the identifying means.

[0790] (Claim 3)

[0791] The system according to claim 1, wherein the presentation means provides personalized feedback based on the user's emotional state.

[0792] "Application example 2 when combining with an emotional engine"

[0793] (Claim 1)

[0794] An analysis means installed on the user's communication device to analyze the visual display,

[0795] An identification means that identifies the challenges faced by the user using the information collected by the analysis means,

[0796] A means of presenting solutions to users based on the problems identified by the specified means,

[0797] An emotion recognition means that recognizes the user's emotional state and adjusts the solution according to that emotion,

[0798] A system that includes this.

[0799] (Claim 2)

[0800] Furthermore, it includes monitoring means to monitor user actions and continuously collect data.

[0801] The system according to claim 1, wherein the identification means uses data acquired by the monitoring means and emotional states acquired by the emotion recognition means.

[0802] (Claim 3)

[0803] The system according to claim 1, comprising a means for providing a solution as audio using voice output technology, and a means for providing guidance information to a user while they are engaged in activities within a commercial facility. [Explanation of symbols]

[0804] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. An analysis means installed on the user's terminal and used to analyze visual information displayed on the screen, A means for identifying problems faced by users using data obtained by the said analysis means, A means of presenting solutions to the user based on the problem identified by the specified means, A system that includes this.

2. Furthermore, it includes monitoring means to monitor user actions and continuously collect data. The system according to claim 1, wherein the data acquired by the monitoring means is used by the identifying means.

3. The system according to claim 1, wherein the presentation means provides the solution as voice using speech synthesis technology.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A