system

A self-service system using facial recognition and voice/text inputs for secure and efficient 24/7 procedure completion addresses the limitations of conventional systems by enabling safe and smooth user interactions.

JP2026104358APending Publication Date: 2026-06-25SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-13
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Conventional systems struggle with handling customer procedures outside business hours, particularly for night shift workers and those with reversed day-night lifestyles, and lack safe personal information handling and quick support for user uncertainties during procedures.

Method used

A self-service system utilizing facial recognition for personal authentication, voice or text-based procedure selection, automatic document generation, and an FAQ database for immediate answers, enabling 24/7 operation and secure procedure completion.

Benefits of technology

Facial recognition ensures secure authentication, while voice and text interfaces guide users through procedures efficiently, allowing safe and smooth completion of tasks anytime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026104358000001_ABST
    Figure 2026104358000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】 Means for performing personal authentication based on face authentication, Means for accepting user's voice or text-based procedure selection, Means for automatically generating documents required for procedures, Means for providing a user interface that displays the generated documents and enables editing, Means for generating an answer using an information database for user questions, Means for confirming completion of procedures, Means for registering or updating payment information, Means for obtaining an answer using a natural language generation engine for questions, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In conventional procedure systems, it has been difficult to handle customers outside business hours, and there has been a lack of self-service available especially for night shift workers and those with a reversed day-night lifestyle. Also, there has been a problem that it is difficult to safely handle personal information in procedures and that quick support cannot be provided for users' uncertainties during the procedure.

Means for Solving the Problems

[0005] This invention provides a means of personal authentication using facial recognition and means of enabling procedure selection by voice or text. Furthermore, it provides means of automatically generating and displaying the necessary documents for the procedure to the user, making the procedure available 24 hours a day. In addition, the generated documents are editable, and answers to user questions are generated immediately using an FAQ database. This realizes a self-service system that enables safe and smooth completion of procedures.

[0006] "Facial recognition" is a technology that verifies an individual's identity based on a user's facial image acquired using a camera.

[0007] "Personal authentication" is the process of verifying that a specific person is who they claim to be.

[0008] "Voice input" is a technology that converts sound into electrical signals and transmits them as information to a computer.

[0009] "Text input" refers to the operation of entering text information into a computer using a keyboard or touch panel.

[0010] "Procedure selection" refers to the act of a user choosing their preferred procedure from a set of available options.

[0011] "Automatic document generation" refers to the process where a program automatically creates documents based on the necessary information.

[0012] An "FAQ database" is a collection of information that includes frequently asked questions and their answers.

[0013] "Procedure completion confirmation" is the process of confirming that all prescribed procedures have been completed without any problems. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0020] In the following embodiments, a labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention relates to a system for users to perform self-service procedures using a terminal. The system includes facial recognition for personal authentication, procedure selection using voice and text, automatic document generation, and a question-answering function utilizing an FAQ database. It also features voice and text interfaces to guide the user through the procedure.

[0036] System Operation Overview

[0037] 1. The user faces the device's camera, and facial recognition is performed. This authenticates the user's identity.

[0038] 2. After successful authentication, the device displays a list of available procedures to the user, allowing them to select a procedure via voice or text.

[0039] 3. The server automatically generates the necessary documents based on the procedure selected by the user and sends them to the terminal. A template containing the information to be filled in is used for this purpose.

[0040] 4. The terminal displays documents to the user and provides an interface that allows the user to review or modify the content.

[0041] 5. If the user asks a question during the process, the terminal will convert the question into text and send it to the server.

[0042] 6. The server consults the FAQ database, generates an appropriate answer to the question, and sends it back to the terminal.

[0043] 7. After the document verification is complete, the user selects "Complete Procedure," at which point the server records the completion of the procedure and sends a confirmation notification to the device.

[0044] Specific example

[0045] For example, if a user wants to change the address on their bank account, they first perform facial recognition on the terminal. Then, they select "Change Address" from the menu, and the system automatically generates the necessary documents for the address change. The user enters the required information into the address change form and confirms it. If the user asks "What documents are required for this?" at some point, the system will respond "You will need identification documents and a certificate of residence." Once the documents have been verified and the user has completed the procedure, the system correctly records the procedure and displays a confirmation message to the user. This system allows users to complete the procedure safely and securely 24 hours a day.

[0046] The following describes the processing flow.

[0047] Step 1:

[0048] The user approaches the device and faces the camera for facial recognition. The device uses the camera to capture an image of the user's face.

[0049] Step 2:

[0050] The device sends the acquired facial image to the server and initiates facial recognition. The server uses the facial recognition system to perform personal authentication and returns the result to the device.

[0051] Step 3:

[0052] The device notifies the user that facial recognition was successful and displays a menu of options. The user selects their desired procedure using the device's display or voice input.

[0053] Step 4:

[0054] When the user selects a procedure, the terminal sends that information to the server. The server selects the necessary document templates for the chosen procedure and automatically generates the documents.

[0055] Step 5:

[0056] The server sends the generated document to the terminal. The terminal displays the document to the user and provides a screen where the entered information can be reviewed and modified.

[0057] Step 6:

[0058] If a user asks a question during the process, the terminal converts the voice input into text and sends the question to the server. The server consults the FAQ database, generates an appropriate answer to the question, and returns it to the terminal.

[0059] Step 7:

[0060] The user reviews the documents and selects "Complete Procedure." The device then sends this instruction to the server.

[0061] Step 8:

[0062] The server confirms that the procedure is complete and sends an acknowledgment notification to the terminal. The terminal displays a confirmation message to the user, informing them that the procedure is finished.

[0063] (Example 1)

[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0065] There is a need for self-service systems that allow users to complete procedures at their own pace, without being restricted by time or location. However, conventional systems lack sufficient features for personal authentication and smooth procedures, resulting in a lack of means for users to complete procedures themselves with peace of mind. In particular, there is a lack of mechanisms to respond quickly and appropriately to user questions. Furthermore, delays in responding to questions that arise during the procedure and the burden of complex operations on users are also problems.

[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0067] In this invention, the server includes means for personal authentication based on facial recognition, means for referencing an information database using an AI model, and means for dynamically guiding the user's procedure progress. This enables the user to proceed with the procedure quickly and efficiently. Specifically, by introducing a secure authentication method using facial recognition and providing consistent support from procedure selection to document generation and answering questions, a safe and reliable 24 / 7 self-service environment is realized.

[0068] "Facial recognition" is an authentication method that analyzes facial features to identify an individual and compares them with registered information in a database.

[0069] "Personal authentication" is a process of verifying that a user is who they claim to be, and it is a technology that ensures security by using specific biometric or identification information.

[0070] "Automatic document generation" is a function that automatically inserts necessary information based on a specific template and generates a document in a predetermined format.

[0071] An "information processing device" is a general term for electronic devices used for inputting, processing, and outputting data, and is sometimes commonly referred to as a computer.

[0072] An "information database" is a collection of data in which various types of data are stored in a structured manner, and information can be searched and analyzed through queries.

[0073] An "AI model" is a mathematical framework for performing a specific task using artificial intelligence technology, and its accuracy is improved through learning algorithms.

[0074] A "prompt statement" is an instruction statement that a user enters into the system to perform an intended operation or obtain information, and is used to trigger a specific process.

[0075] One embodiment of the present invention is a system that enables users to perform procedures efficiently and securely through self-service. This system includes facial recognition, procedure selection, automatic document generation, question answering, and progress guidance functions.

[0076] The user first authenticates themselves using facial recognition via the device's camera. The facial recognition technology employs AI-powered facial feature extraction and matching algorithms. This ensures the protection of the user's personal information and provides secure authentication.

[0077] After successful user authentication, the terminal displays a list of available procedures on its menu screen. The user selects their desired procedure using voice input or touch screen operation. This allows the user to easily choose the procedure they need.

[0078] The server automatically generates the necessary documents according to the user's selected procedure. These documents are generated based on pre-configured templates, and by inserting the user's information, the document creation process is expedited.

[0079] Furthermore, the system can utilize a generative AI model to refer to an FAQ database and provide appropriate answers to user questions. In this process, the AI ​​analyzes the question, quickly searches for relevant data, and generates the answer. Natural language processing technology is used in this process.

[0080] Furthermore, throughout the process, the terminal guides the user through voice and text interfaces, helping them to operate safely and intuitively.

[0081] For example, if a user wants to change the address on their bank account, a possible prompt message might be, "Please tell me what documents are required to change the address on my bank account." The system responds to this prompt by providing the necessary documents and procedural details, supporting the user in completing the process smoothly.

[0082] As described above, the present invention enables users to proceed with procedures with peace of mind and efficiency.

[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0084] Step 1:

[0085] The user provides facial recognition input data by facing their face towards the device's camera. The device sends the acquired facial image data to an AI-powered facial recognition engine and performs data processing to match it with personal identification information. If the match is successful, it generates an output indicating that personal authentication is complete and proceeds to the next step.

[0086] Step 2:

[0087] The terminal displays multiple procedure options on the screen to users who have successfully authenticated. This display includes information about available procedures based on the system's database. Users make their selection via voice input or touch operation. The selected procedure information is entered, and an output indicating that the procedure selection is complete is generated.

[0088] Step 3:

[0089] The server selects the necessary document template based on the user's procedure selection, inserts user data, and performs automated document generation. During this process, a template engine is used to process the specified information. The generated document is output to the terminal and provided to the user.

[0090] Step 4:

[0091] The terminal displays the generated document on its screen and provides an interface that allows the user to review and edit the document content. After the user proofreads the document and enters the necessary information, they review the results on the terminal. The terminal generates output indicating that the review is complete and proceeds to the next step.

[0092] Step 5:

[0093] If the user enters a question during the process, the terminal receives the voice or text input and prepares to send the question to the server. This question is treated as a prompt and used as input data to the server.

[0094] Step 6:

[0095] Based on the received prompt, the server utilizes a generative AI model to refer to the FAQ database. This database search performs data calculations to generate the optimal answer to the question. The server then generates output in text format, which is sent to the terminal.

[0096] Step 7:

[0097] When the user chooses to complete the procedure, the server records the completion status and notifies the terminal accordingly. This process inputs a completion notification, and the user receives output that allows them to confirm the completion of the procedure.

[0098] (Application Example 1)

[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0100] In modern online procedures, the complexity of the process is a problem because users are required to take multiple steps. In particular, when registering or updating payment information, the progress of the procedure is often unclear, and it can be difficult to understand what documents and information are needed. There are also challenges in obtaining accurate and prompt answers to user questions. A system that can improve the user experience by solving these problems is needed.

[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0102] In this invention, the server includes means for performing personal authentication based on facial recognition, means for receiving the user's procedure selection by voice or text, means for registering or updating payment information, means for obtaining answers to questions using a natural language generation engine with respect to an information database, and means having a voice and text interface for guiding the procedure progress and payment operations. This enables the automation and efficiency of procedures, allowing users to perform procedures that can be accepted simply and quickly.

[0103] "Facial recognition" is a technology that identifies individuals by analyzing the features of a user's face as digital data.

[0104] "Personal authentication" is a process for verifying that a specific user is who they claim to be.

[0105] "Procedure selection by voice or text" is a feature that allows users to select the procedure they wish to perform through voice input or text input.

[0106] "Methods for automatically generating necessary documents for procedures" refers to technologies that automatically create the necessary procedural documents based on pre-configured templates and information provided by the user.

[0107] A "user interface" is a screen or control panel that provides an efficient and convenient means for a system and a user to exchange information.

[0108] An "information database" is a collection of information used to provide answers to questions.

[0109] "Procedure completion confirmation" is a function that notifies and records to the user that all steps of the selected procedure have been successfully completed.

[0110] "Means for registering or updating payment information" refers to the process of registering a user's new payment information in the system and changing it as needed.

[0111] A "natural language generation engine" is a technology that automatically generates human-readable text in response to input information or questions.

[0112] A "voice and text interface" is a common gateway for users and systems to communicate via voice commands and text messages.

[0113] To implement this invention, a system is needed in which both the server and the user terminal cooperate to enable the user to perform procedures efficiently. First, the server uses facial recognition technology to authenticate the user's identity. This facial recognition process uses an image processing library to acquire facial images in real time using the user terminal's camera. Next, the server uses speech recognition technology to provide an interface for the user to select procedures by voice or text. Here, speech-to-text conversion technology plays a crucial role.

[0114] When a user selects a specific procedure, the server automatically generates the necessary procedural documents using a template engine and sends them to the user's terminal. The user is then provided with an interface on their terminal that allows them to view the documents and edit the information as needed. This interface is designed to be intuitive and easy to use.

[0115] If a user asks a question during the process, the server uses a generative AI model to retrieve the appropriate answer from the information database and provides the user with a response generated in natural language. For example, if a user asks, "What documents are required for this procedure?", the server will respond, "You will need identification documents."

[0116] Furthermore, once the process is complete, the server will inform the user of its completion and provide guidance on how to proceed to the next step. Through this entire process, users can register or update their payment information with confidence.

[0117] An example of a prompt message is, "What information is required when registering a new payment method?", and an accurate answer is generated based on the information database.

[0118] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0119] Step 1:

[0120] The server performs facial recognition of the user. The user faces the camera on their device, and the device captures the facial image and sends it to the server. The server analyzes the received image using an image processing library and compares it with registered facial data to identify the individual. The input is the camera image, and the output is the authentication result.

[0121] Step 2:

[0122] After successful authentication, the server sends a list of available procedures to the terminal. The terminal displays this list and prompts the user to select a procedure via voice or text. The input is the authentication result, and the output is a display of the procedure list.

[0123] Step 3:

[0124] When a user selects a procedure by voice, the terminal converts the voice input into text and sends it to the server. The server analyzes the transmitted text to identify the selected procedure. The input is voice data, and the output is text information of the selected procedure.

[0125] Step 4:

[0126] Based on the specified procedure, the server creates the necessary documents using an automatically generated template. The generated documents are sent to the terminal, which displays them to the user. The user can then review and edit the document contents. The input is the selected procedure information, and the output is the generated document data.

[0127] Step 5:

[0128] If a user asks a question during the process, the terminal transcribes the question into text and sends it to the server. The server uses a generative AI model to search a relevant information database, generates an accurate answer to the question, and sends it back to the terminal. The input is the question text, and the output is the generated answer.

[0129] Step 6:

[0130] When the user selects to complete the procedure, the server verifies that all processes have terminated properly and sends the result to the terminal. The terminal then displays a completion notification to the user. The input is the procedure completion instruction, and the output is the completion notification.

[0131] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0132] This invention is a self-service system incorporating an emotion engine, which aims to automate various procedures and improve convenience for users using a terminal. The emotion engine analyzes the user's facial expressions and voice to recognize their emotional state and utilizes that information for procedural guidance and response generation.

[0133] System Operation Overview

[0134] 1. The user approaches the device and performs facial recognition. The device captures the facial image and sends it to the server for authentication.

[0135] 2. Based on the authentication result, the server generates a menu of procedures that the user can select and sends it to the terminal.

[0136] 3. The terminal displays a procedure menu to the user and provides the ability to select a procedure by voice or text. At this point, the emotion engine operates, analyzing the user's facial expressions and tone of voice to recognize their emotions.

[0137] 4. Recognized emotional information is reflected in the content and tone of the procedural guide. If the user is experiencing stress, the guide will be adjusted to provide a more gentle approach.

[0138] 5. Based on the user's selection, the server automatically generates the necessary documents and sends them to the terminal. The terminal displays the documents to the user and provides an editable interface.

[0139] 6. When a question arises, the user's voice is transcribed into text and sent to the server. An appropriate answer is generated by referring to the FAQ database. Here again, sentiment information is used to adjust the tone and detail of the answer.

[0140] 7. Once the procedure is complete, the server will perform a final check and send and display an approval message to the terminal.

[0141] Specific example

[0142] For example, when a user is going through the process of purchasing travel insurance, if the terminal detects an anxious expression on their face, the system will guide them through the process in a calming tone. It will then generate a helpful inquiry such as, "Is there anything I can help you with?" to support the user. If the user asks, "Does this insurance cover pre-existing conditions?", the system will provide a detailed answer from the FAQ, adjusting the tone of the explanation according to the user's emotions. In this way, a user-friendly self-service can be provided.

[0143] The following describes the processing flow.

[0144] Step 1:

[0145] The user faces the device, initiating facial recognition. The device captures an image of the user's face using its camera and sends the image data to the server for personal authentication.

[0146] Step 2:

[0147] The server processes the received facial image data using its facial recognition system and sends the authentication result back to the terminal. If authentication is successful, the process proceeds to the next step.

[0148] Step 3:

[0149] The terminal notifies the user of successful authentication and displays a procedure menu. The user selects their desired procedure using voice or the touch panel. The emotion engine analyzes the user's facial expressions and voice to determine their emotions.

[0150] Step 4:

[0151] The server identifies the necessary document templates based on the user's selected procedure and automatically generates and sends the documents to the terminal. The content and tone of the procedure guidance are adjusted based on the recognized emotions.

[0152] Step 5:

[0153] The terminal provides the user with an interface for viewing, editing, and reviewing documents. If the user experiences anxiety or questions during the process, the emotion engine detects this and prepares to provide appropriate guidance.

[0154] Step 6:

[0155] When a user asks a question, the device converts the voice input into text and sends it to the server. The server searches the FAQ database and generates the most relevant answer. The tone and detail of the answer are adjusted according to the user's mood.

[0156] Step 7:

[0157] Once the user completes the procedure, the device sends that information to the server. The server confirms that the procedure was completed successfully and sends a final confirmation message to the device to notify the user.

[0158] Step 8:

[0159] The device displays a final confirmation message to inform the user that the procedure is complete. After confirmation, the user can leave the device.

[0160] (Example 2)

[0161] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0162] Current self-service systems lack consideration for users' emotional states, making it difficult for users to perform procedures with peace of mind. Furthermore, the lack of dynamic adjustments to the procedure flow based on individual user circumstances leads to decreased operational efficiency and user satisfaction.

[0163] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0164] In this invention, the server includes means for identifying an individual based on facial recognition data, means for accepting process selections based on the user's speech or input information, and means for analyzing the user's emotional state and adjusting the guidance content. This enables responses that take the user's emotions into consideration and provides a flexible procedural flow tailored to individual situations.

[0165] "Facial recognition data" refers to information extracted from facial images taken to identify individual characteristics.

[0166] "Means of identifying individuals" refers to processes and technologies that use facial recognition data to distinguish a specific individual from others.

[0167] "User utterances or input information" refers to voice or text data used by users to communicate their choices or intentions to the system.

[0168] A "means for accepting process selection" refers to a mechanism for receiving input related to the user's procedure selection and initiating processing based on that input.

[0169] An "electronic document" is a digital document containing information necessary for a process, which can be created, edited, and displayed on a computer.

[0170] A "knowledge database" is a collection of data that stores information and answers in a specific field, and is used to respond appropriately to questions from users.

[0171] "Means of analyzing emotional states" refers to technical methods that analyze a user's facial expressions and voice to determine their emotions.

[0172] "Means for adjusting guidance content" refers to a mechanism for changing the content and tone of guidance provided to the user based on their analyzed emotional state.

[0173] This invention provides a self-service system incorporating an emotion engine, allowing users to automate various procedures and improve convenience through their user terminals. The system performs personal identification based on user facial recognition data and analyzes their emotional state to provide flexible guidance tailored to the user's situation.

[0174] Specifically, when a user approaches the terminal, the facial recognition system activates and captures an image of the user's face. The hardware used includes a high-performance camera, and the software includes image processing libraries such as OpenCV. The terminal sends the captured facial image data to the server for personal identification processing. Machine learning algorithms are used in this identification process. If facial recognition is successful, the server creates a procedural menu related to the user and sends it to the terminal.

[0175] Users can select their desired task from a procedure menu displayed on the terminal using voice or text input. In this step, an emotion engine analyzes the user's facial expressions and voice tone to evaluate the user's emotional state. This analysis uses a Python®-based library, such as the Emotion Recognition API. The tone and content of the guidance are adjusted according to the evaluated emotional state. For example, if the user is feeling anxious, the terminal will provide a friendly voice guidance to calm the user.

[0176] Once a procedure is selected, the server automatically generates the necessary electronic documents using document templates. These generated electronic documents are displayed on the terminal in an editable format. If the user asks additional questions, these questions are transcribed using speech recognition technology, and answers are generated by referencing the server's knowledge database. This ensures that an appropriate response is selected, taking the user's feelings into consideration.

[0177] For example, when a user is processing travel insurance on a terminal, if the emotion engine detects the user's anxiety, it will add calming phrases such as "Is there anything I can help you with?" in its guidance. Next, if the user asks a question, either by typing or speaking, such as "Please tell me what is covered by this insurance," the system will provide a detailed answer from its knowledge database.

[0178] An example of a prompt message could be: "Create a sample conversation that provides a travel insurance procedure guide based on the user's emotions. Include examples of guidance tailored to different emotional states."

[0179] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0180] Step 1:

[0181] When a user approaches the device, the device uses its camera to capture the user's face. This facial image data becomes the input, and the device sends this data to the server. The server processes the received image data through a facial recognition algorithm and outputs the user's personal identification information. This process involves data analysis using an image processing library.

[0182] Step 2:

[0183] The server generates a menu of procedures that the user can execute based on the facial recognition results. The input here is the recognized personal identification information. The server refers to the user's attribute information and history to dynamically determine the appropriate procedure menu and outputs that menu to the terminal in JSON format.

[0184] Step 3:

[0185] The terminal displays the procedure menu received from the server on its user interface. The user makes selections using voice or touch, based on this displayed menu. The terminal sends the selected information to the server, triggering the progress of the corresponding procedure.

[0186] Step 4:

[0187] Once the user's procedure selection is transmitted to the server, the server activates the emotion engine and analyzes the user's emotional state. The input for the analysis is the user's facial expressions and voice data. The server passes this data through the emotion recognition system to detect the user's emotional state and outputs the result to the terminal as emotional information.

[0188] Step 5:

[0189] The server adjusts the tone and content of the procedural guidance based on the user's emotional information. Input consists of information necessary for the procedure to progress and emotional data. Based on this information, the server generates reassuring guidance content and outputs it to the terminal.

[0190] Step 6:

[0191] Based on the procedure selected by the user, the server automatically generates the necessary electronic documents from document templates, following prompts. The input is the selected procedure information, and the output is an editable electronic document. This document is sent to the terminal and displayed to the user. The user can edit the electronic document as needed.

[0192] Step 7:

[0193] During the process, if a user asks a question, the audio is converted to text by the device and sent to the server. The server consults a knowledge database, generates an appropriate answer using the transcribed question as input, adjusts the emotional tone, and then outputs the answer to the device.

[0194] Step 8:

[0195] When the process reaches its final stage, the server verifies that all steps have been completed correctly. After verification, the server generates an approval message and sends it to the terminal. The terminal displays this message to the user, informing them that the process has been successfully completed.

[0196] (Application Example 2)

[0197] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0198] In modern self-service systems, information is often provided uniformly without regard for the user's feelings, which frequently leads to confusion and stress. To address this problem, it is necessary to recognize the user's emotional state in real time and provide personalized guidance and answers accordingly, thereby creating a smoother and more effective self-service experience.

[0199] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0200] In this invention, the server includes a function to recognize the user's emotions and adjust the procedure guide and responses accordingly, a function to perform individual authentication based on facial recognition, and a function to accept the user's procedure selection via voice or text. This makes it possible to provide customized procedure guidance and responses according to the user's emotional state.

[0201] "Facial recognition" is a technology that analyzes the facial features of a user to identify them as an individual.

[0202] "Individual authentication" is a procedure used by a system to identify a specific user and grant them access.

[0203] "User" refers to the end user who operates this system.

[0204] "Voice or text" refers to a method of communication used by users to input procedures or questions into a system.

[0205] The "function to accept procedure selection" refers to the ability to provide an interface for users to choose the procedure they wish to perform.

[0206] The "automatic document generation function" is a technology that automatically creates the documents necessary for a procedure based on an algorithm.

[0207] An "FAQ database" is a source of information that compiles frequently asked questions and their answers.

[0208] The "function that recognizes emotions and adjusts procedural guidance and responses accordingly" is a technology that analyzes the user's emotional state and dynamically changes the guidance method and content based on that analysis.

[0209] An "information processing device" is a general term for a series of electronic devices that collect, analyze, and output data.

[0210] The system for realizing this invention is a self-service device that recognizes the user's emotions and adjusts the operation guide and responses accordingly. The system uses a facial recognition camera to capture the user's face when the user approaches the terminal and authenticates a specific individual. This is achieved by utilizing common facial recognition technologies (e.g., OpenCV or the dlib library).

[0211] Subsequently, the user selects the necessary procedure via voice or text input, and the system generates and displays the necessary information for that procedure. This information generation utilizes natural language processing libraries (e.g., spaCy or NLTK). If the user asks a question during the procedure, the voice is converted into text, and the system searches the server's FAQ database for the most appropriate answer and provides it. The FAQ database is often constructed using an SQL database.

[0212] Furthermore, an emotion recognition engine (e.g., Microsoft® Azure® Emotion API or Google® Cloud Vision API) analyzes the user's facial expressions and voice tone, and the guidance method and response tone are adjusted accordingly. This system allows users to use self-services efficiently and stress-free.

[0213] For example, if a customer is using a self-checkout machine in a physical store to purchase a new product and the terminal detects their confused expression, the system will display a helpful message such as, "Is this your first time using this? Do you need assistance?" Additional guidance would also be provided, including detailed instructions on how to use the product and its features.

[0214] An example of a prompt message is, "How can I adjust and display the procedure guide based on the emotions a customer feels when using a self-checkout?"

[0215] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0216] Step 1:

[0217] The device detects the user and captures a facial image using a facial recognition camera. The input is image data acquired by the camera, and the output is the identification information of the recognized user. User authentication is performed using this data, and the individual is identified.

[0218] Step 2:

[0219] The server generates a menu of procedures that the user can select from, based on the authentication result. The input is the user's identification information, and the output is a list of procedure options. The server retrieves the user's history data from the database and presents the most appropriate menu.

[0220] Step 3:

[0221] The user selects their desired procedure on the terminal via voice or text input. The input is the user's voice or text data, and the output is the content of the selected procedure. Voice recognition technology is used to convert the voice into text and analyze the user's selection.

[0222] Step 4:

[0223] The server automatically generates the necessary documents according to the selected procedure and sends them to the terminal. The input is information about the selected procedure, and the output is the generated document data. Natural language processing is used to organize and provide the documents in a way that is easy for the user to understand.

[0224] Step 5:

[0225] The terminal displays documents to the user and provides an editable interface. Input is document data from the server, and output is the user's edited document. The system allows the user to modify the document as needed.

[0226] Step 6:

[0227] When a user asks a question, the audio is converted into text by the device and sent to the server. The input is the user's question in audio form, and the output is the converted question text. An appropriate answer is then selected from the FAQ database.

[0228] Step 7:

[0229] The server analyzes the user's emotional state through an emotion recognition engine, adjusts the tone of its response based on that analysis, and sends the response to the terminal. Input is the user's facial expressions and voice data, and output is the adjusted response. The system communicates flexibly according to the user's emotions.

[0230] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0231] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0232] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0233] [Second Embodiment]

[0234] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0235] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0236] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0237] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0238] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0239] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0240] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0241] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0242] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0243] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0244] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0245] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0246] This invention relates to a system for users to perform self-service procedures using a terminal. The system includes facial recognition for personal authentication, procedure selection using voice and text, automatic document generation, and a question-answering function utilizing an FAQ database. It also features voice and text interfaces to guide the user through the procedure.

[0247] System Operation Overview

[0248] 1. The user faces the device's camera, and facial recognition is performed. This authenticates the user's identity.

[0249] 2. After successful authentication, the device displays a list of available procedures to the user, allowing them to select a procedure via voice or text.

[0250] 3. The server automatically generates the necessary documents based on the procedure selected by the user and sends them to the terminal. A template containing the information to be filled in is used for this purpose.

[0251] 4. The terminal displays documents to the user and provides an interface that allows the user to review or modify the content.

[0252] 5. If the user asks a question during the process, the terminal will convert the question into text and send it to the server.

[0253] 6. The server consults the FAQ database, generates an appropriate answer to the question, and sends it back to the terminal.

[0254] 7. After the document verification is complete, the user selects "Complete Procedure," at which point the server records the completion of the procedure and sends a confirmation notification to the device.

[0255] Specific example

[0256] For example, if a user wants to change the address on their bank account, they first perform facial recognition on the terminal. Then, they select "Change Address" from the menu, and the system automatically generates the necessary documents for the address change. The user enters the required information into the address change form and confirms it. If the user asks "What documents are required for this?" at some point, the system will respond "You will need identification documents and a certificate of residence." Once the documents have been verified and the user has completed the procedure, the system correctly records the procedure and displays a confirmation message to the user. This system allows users to complete the procedure safely and securely 24 hours a day.

[0257] The following describes the processing flow.

[0258] Step 1:

[0259] The user approaches the device and faces the camera for facial recognition. The device uses the camera to capture an image of the user's face.

[0260] Step 2:

[0261] The device sends the acquired facial image to the server and initiates facial recognition. The server uses the facial recognition system to perform personal authentication and returns the result to the device.

[0262] Step 3:

[0263] The device notifies the user that facial recognition was successful and displays a menu of options. The user selects their desired procedure using the device's display or voice input.

[0264] Step 4:

[0265] When the user selects a procedure, the terminal sends that information to the server. The server selects the necessary document templates for the chosen procedure and automatically generates the documents.

[0266] Step 5:

[0267] The server sends the generated document to the terminal. The terminal displays the document to the user and provides a screen where the entered information can be reviewed and modified.

[0268] Step 6:

[0269] If a user asks a question during the process, the terminal converts the voice input into text and sends the question to the server. The server consults the FAQ database, generates an appropriate answer to the question, and returns it to the terminal.

[0270] Step 7:

[0271] The user reviews the documents and selects "Complete Procedure." The device then sends this instruction to the server.

[0272] Step 8:

[0273] The server confirms that the procedure is complete and sends an acknowledgment notification to the terminal. The terminal displays a confirmation message to the user, informing them that the procedure is finished.

[0274] (Example 1)

[0275] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0276] There is a need for self-service systems that allow users to complete procedures at their own pace, without being restricted by time or location. However, conventional systems lack sufficient features for personal authentication and smooth procedures, resulting in a lack of means for users to complete procedures themselves with peace of mind. In particular, there is a lack of mechanisms to respond quickly and appropriately to user questions. Furthermore, delays in responding to questions that arise during the procedure and the burden of complex operations on users are also problems.

[0277] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0278] In this invention, the server includes means for personal authentication based on facial recognition, means for referencing an information database using an AI model, and means for dynamically guiding the user's procedure progress. This enables the user to proceed with the procedure quickly and efficiently. Specifically, by introducing a secure authentication method using facial recognition and providing consistent support from procedure selection to document generation and answering questions, a safe and reliable 24 / 7 self-service environment is realized.

[0279] "Facial recognition" is an authentication method that analyzes facial features to identify an individual and compares them with registered information in a database.

[0280] "Personal authentication" is a process of verifying that the user is the actual person, and it is a technology that ensures security by using specific biometric information or identification information.

[0281] "Document automatic generation" is a function that automatically inserts necessary information based on a specific template and generates a document in a predetermined format.

[0282] "Information processing device" is a general term for electronic devices that perform data input, processing, and output, and is sometimes generally called a computer.

[0283] "Information database" is a data set in which various data are structured and stored, and information search and analysis are performed through queries.

[0284] "AI model" is a mathematical framework for executing specific tasks using artificial intelligence technology, and aims to improve accuracy through learning algorithms.

[0285] "Prompt sentence" is an instruction sentence input by the user to the system in order to perform the intended operation or obtain information, and is a sentence used to induce specific processing.

[0286] The embodiment for implementing the present invention is a system for users to perform procedures efficiently and securely in a self-service manner. This system includes face authentication, procedure selection, document automatic generation, question and answer, and progress guidance functions.

[0287] The user first uses the camera of the terminal to perform face authentication to authenticate himself. AI-based face feature extraction and matching algorithms are used in face authentication technology. Thereby, the user's personal information is protected and secure authentication is realized.

[0288] After successful user authentication, the terminal displays a list of available procedures on its menu screen. The user selects their desired procedure using voice input or touch screen operation. This allows the user to easily choose the procedure they need.

[0289] The server automatically generates the necessary documents according to the user's selected procedure. These documents are generated based on pre-configured templates, and by inserting the user's information, the document creation process is expedited.

[0290] Furthermore, the system can utilize a generative AI model to refer to an FAQ database and provide appropriate answers to user questions. In this process, the AI ​​analyzes the question, quickly searches for relevant data, and generates the answer. Natural language processing technology is used in this process.

[0291] Furthermore, throughout the process, the terminal guides the user through voice and text interfaces, helping them to operate safely and intuitively.

[0292] For example, if a user wants to change the address on their bank account, a possible prompt message might be, "Please tell me what documents are required to change the address on my bank account." The system responds to this prompt by providing the necessary documents and procedural details, supporting the user in completing the process smoothly.

[0293] As described above, the present invention enables users to proceed with procedures with peace of mind and efficiency.

[0294] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0295] Step 1:

[0296] The user provides facial recognition input data by facing their face towards the device's camera. The device sends the acquired facial image data to an AI-powered facial recognition engine and performs data processing to match it with personal identification information. If the match is successful, it generates an output indicating that personal authentication is complete and proceeds to the next step.

[0297] Step 2:

[0298] The terminal displays multiple procedure options on the screen to users who have successfully authenticated. This display includes information about available procedures based on the system's database. Users make their selection via voice input or touch operation. The selected procedure information is entered, and an output indicating that the procedure selection is complete is generated.

[0299] Step 3:

[0300] The server selects the necessary document template based on the user's procedure selection, inserts user data, and performs automated document generation. During this process, a template engine is used to process the specified information. The generated document is output to the terminal and provided to the user.

[0301] Step 4:

[0302] The terminal displays the generated document on its screen and provides an interface that allows the user to review and edit the document content. After the user proofreads the document and enters the necessary information, they review the results on the terminal. The terminal generates output indicating that the review is complete and proceeds to the next step.

[0303] Step 5:

[0304] If the user enters a question during the process, the terminal receives the voice or text input and prepares to send the question to the server. This question is treated as a prompt and used as input data to the server.

[0305] Step 6:

[0306] Based on the received prompt sentence, the server utilizes the generative AI model to refer to the FAQ database. Through this database search, data operations are performed to generate the optimal answer to the question. An appropriate answer is obtained in text form, and an output to be sent to the terminal is generated.

[0307] Step 7:

[0308] When the user selects to complete the procedure, the server records the completion status of the procedure and notifies the terminal to that effect. Through this process, a completion notification is input, and an output is generated that allows the user to confirm the end of the procedure.

[0309] (Application Example 1)

[0310] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0311] In modern online procedures, users require multiple steps, making the complexity of the procedure a problem. Especially when registering or updating payment information, etc., the progress of the procedure may not be clear, and it may be difficult to understand the necessary documents and information. There are also problems such as not being able to smoothly obtain an accurate answer to the user's question. There is a need for a system that can improve the user experience by solving these problems.

[0312] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.

[0313] In this invention, the server includes means for performing personal authentication based on facial recognition, means for receiving the user's procedure selection by voice or text, means for registering or updating payment information, means for obtaining answers to questions using a natural language generation engine with respect to an information database, and means having a voice and text interface for guiding the procedure progress and payment operations. This enables the automation and efficiency of procedures, allowing users to perform procedures that can be accepted simply and quickly.

[0314] "Facial recognition" is a technology that identifies individuals by analyzing the features of a user's face as digital data.

[0315] "Personal authentication" is a process for verifying that a specific user is who they claim to be.

[0316] "Procedure selection by voice or text" is a feature that allows users to select the procedure they wish to perform through voice input or text input.

[0317] "Methods for automatically generating necessary documents for procedures" refers to technologies that automatically create the necessary procedural documents based on pre-configured templates and information provided by the user.

[0318] A "user interface" is a screen or control panel that provides an efficient and convenient means for a system and a user to exchange information.

[0319] An "information database" is a collection of information used to provide answers to questions.

[0320] "Procedure completion confirmation" is a function that notifies and records to the user that all steps of the selected procedure have been successfully completed.

[0321] "Means for registering or updating payment information" refers to the process of registering a user's new payment information in the system and changing it as needed.

[0322] A "natural language generation engine" is a technology that automatically generates human-readable text in response to input information or questions.

[0323] A "voice and text interface" is a common gateway for users and systems to communicate via voice commands and text messages.

[0324] To implement this invention, a system is needed in which both the server and the user terminal cooperate to enable the user to perform procedures efficiently. First, the server uses facial recognition technology to authenticate the user's identity. This facial recognition process uses an image processing library to acquire facial images in real time using the user terminal's camera. Next, the server uses speech recognition technology to provide an interface for the user to select procedures by voice or text. Here, speech-to-text conversion technology plays a crucial role.

[0325] When a user selects a specific procedure, the server automatically generates the necessary procedural documents using a template engine and sends them to the user's terminal. The user is then provided with an interface on their terminal that allows them to view the documents and edit the information as needed. This interface is designed to be intuitive and easy to use.

[0326] If a user asks a question during the process, the server uses a generative AI model to retrieve the appropriate answer from the information database and provides the user with a response generated in natural language. For example, if a user asks, "What documents are required for this procedure?", the server will respond, "You will need identification documents."

[0327] Furthermore, once the process is complete, the server will inform the user of its completion and provide guidance on how to proceed to the next step. Through this entire process, users can register or update their payment information with confidence.

[0328] An example of a prompt message is, "What information is required when registering a new payment method?", and an accurate answer is generated based on the information database.

[0329] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0330] Step 1:

[0331] The server performs facial recognition of the user. The user faces the camera on their device, and the device captures the facial image and sends it to the server. The server analyzes the received image using an image processing library and compares it with registered facial data to identify the individual. The input is the camera image, and the output is the authentication result.

[0332] Step 2:

[0333] After successful authentication, the server sends a list of available procedures to the terminal. The terminal displays this list and prompts the user to select a procedure via voice or text. The input is the authentication result, and the output is a display of the procedure list.

[0334] Step 3:

[0335] When a user selects a procedure by voice, the terminal converts the voice input into text and sends it to the server. The server analyzes the transmitted text to identify the selected procedure. The input is voice data, and the output is text information of the selected procedure.

[0336] Step 4:

[0337] Based on the specified procedure, the server creates the necessary documents using an automatically generated template. The generated documents are sent to the terminal, which displays them to the user. The user can then review and edit the document contents. The input is the selected procedure information, and the output is the generated document data.

[0338] Step 5:

[0339] If a user asks a question during the process, the terminal transcribes the question into text and sends it to the server. The server uses a generative AI model to search a relevant information database, generates an accurate answer to the question, and sends it back to the terminal. The input is the question text, and the output is the generated answer.

[0340] Step 6:

[0341] When the user selects to complete the procedure, the server verifies that all processes have terminated properly and sends the result to the terminal. The terminal then displays a completion notification to the user. The input is the procedure completion instruction, and the output is the completion notification.

[0342] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0343] This invention is a self-service system incorporating an emotion engine, which aims to automate various procedures and improve convenience for users using a terminal. The emotion engine analyzes the user's facial expressions and voice to recognize their emotional state and utilizes that information for procedural guidance and response generation.

[0344] System Operation Overview

[0345] 1. The user approaches the device and performs facial recognition. The device captures the facial image and sends it to the server for authentication.

[0346] 2. Based on the authentication result, the server generates a menu of procedures that the user can select and sends it to the terminal.

[0347] 3. The terminal displays a procedure menu to the user and provides the ability to select a procedure by voice or text. At this point, the emotion engine operates, analyzing the user's facial expressions and tone of voice to recognize their emotions.

[0348] 4. Recognized emotional information is reflected in the content and tone of the procedural guide. If the user is experiencing stress, the guide will be adjusted to provide a more gentle approach.

[0349] 5. Based on the user's selection, the server automatically generates the necessary documents and sends them to the terminal. The terminal displays the documents to the user and provides an editable interface.

[0350] 6. When a question arises, the user's voice is transcribed into text and sent to the server. An appropriate answer is generated by referring to the FAQ database. Here again, sentiment information is used to adjust the tone and detail of the answer.

[0351] 7. Once the procedure is complete, the server will perform a final check and send and display an approval message to the terminal.

[0352] Specific example

[0353] For example, when a user is going through the process of purchasing travel insurance, if the terminal detects an anxious expression on their face, the system will guide them through the process in a calming tone. It will then generate a helpful inquiry such as, "Is there anything I can help you with?" to support the user. If the user asks, "Does this insurance cover pre-existing conditions?", the system will provide a detailed answer from the FAQ, adjusting the tone of the explanation according to the user's emotions. In this way, a user-friendly self-service can be provided.

[0354] The following describes the processing flow.

[0355] Step 1:

[0356] The user faces the device, initiating facial recognition. The device captures an image of the user's face using its camera and sends the image data to the server for personal authentication.

[0357] Step 2:

[0358] The server processes the received facial image data using its facial recognition system and sends the authentication result back to the terminal. If authentication is successful, the process proceeds to the next step.

[0359] Step 3:

[0360] The terminal notifies the user of successful authentication and displays a procedure menu. The user selects their desired procedure using voice or the touch panel. The emotion engine analyzes the user's facial expressions and voice to determine their emotions.

[0361] Step 4:

[0362] The server identifies the necessary document templates based on the user's selected procedure and automatically generates and sends the documents to the terminal. The content and tone of the procedure guidance are adjusted based on the recognized emotions.

[0363] Step 5:

[0364] The terminal provides the user with an interface for viewing, editing, and reviewing documents. If the user experiences anxiety or questions during the process, the emotion engine detects this and prepares to provide appropriate guidance.

[0365] Step 6:

[0366] When a user asks a question, the device converts the voice input into text and sends it to the server. The server searches the FAQ database and generates the most relevant answer. The tone and detail of the answer are adjusted according to the user's mood.

[0367] Step 7:

[0368] Once the user completes the procedure, the device sends that information to the server. The server confirms that the procedure was completed successfully and sends a final confirmation message to the device to notify the user.

[0369] Step 8:

[0370] The device displays a final confirmation message to inform the user that the procedure is complete. After confirmation, the user can leave the device.

[0371] (Example 2)

[0372] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0373] Current self-service systems lack consideration for users' emotional states, making it difficult for users to perform procedures with peace of mind. Furthermore, the lack of dynamic adjustments to the procedure flow based on individual user circumstances leads to decreased operational efficiency and user satisfaction.

[0374] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0375] In this invention, the server includes means for identifying an individual based on facial recognition data, means for accepting process selections based on the user's speech or input information, and means for analyzing the user's emotional state and adjusting the guidance content. This enables responses that take the user's emotions into consideration and provides a flexible procedural flow tailored to individual situations.

[0376] "Facial recognition data" refers to information extracted from facial images taken to identify individual characteristics.

[0377] "Means of identifying individuals" refers to processes and technologies that use facial recognition data to distinguish a specific individual from others.

[0378] "User utterances or input information" refers to voice or text data used by users to communicate their choices or intentions to the system.

[0379] A "means for accepting process selection" refers to a mechanism for receiving input related to the user's procedure selection and initiating processing based on that input.

[0380] An "electronic document" is a digital document containing information necessary for a process, which can be created, edited, and displayed on a computer.

[0381] A "knowledge database" is a collection of data that stores information and answers in a specific field, and is used to respond appropriately to questions from users.

[0382] "Means of analyzing emotional states" refers to technical methods that analyze a user's facial expressions and voice to determine their emotions.

[0383] "Means for adjusting guidance content" refers to a mechanism for changing the content and tone of guidance provided to the user based on their analyzed emotional state.

[0384] This invention provides a self-service system incorporating an emotion engine, allowing users to automate various procedures and improve convenience through their user terminals. The system performs personal identification based on user facial recognition data and analyzes their emotional state to provide flexible guidance tailored to the user's situation.

[0385] Specifically, when a user approaches the terminal, the facial recognition system activates and captures an image of the user's face. The hardware used includes a high-performance camera, and the software includes image processing libraries such as OpenCV. The terminal sends the captured facial image data to the server for personal identification processing. Machine learning algorithms are used in this identification process. If facial recognition is successful, the server creates a procedural menu related to the user and sends it to the terminal.

[0386] Users can select their desired task from a procedure menu displayed on the terminal using voice or text input. In this step, an emotion engine analyzes the user's facial expressions and voice tone to evaluate the user's emotional state. This analysis uses a Python-based library, such as the Emotion Recognition API. The tone and content of the guidance are adjusted according to the evaluated emotional state. For example, if the user is feeling anxious, the terminal will provide a friendly voice guidance to calm the user.

[0387] Once a procedure is selected, the server automatically generates the necessary electronic documents using document templates. These generated electronic documents are displayed on the terminal in an editable format. If the user asks additional questions, these questions are transcribed using speech recognition technology, and answers are generated by referencing the server's knowledge database. This ensures that an appropriate response is selected, taking the user's feelings into consideration.

[0388] For example, when a user is processing travel insurance on a terminal, if the emotion engine detects the user's anxiety, it will add calming phrases such as "Is there anything I can help you with?" in its guidance. Next, if the user asks a question, either by typing or speaking, such as "Please tell me what is covered by this insurance," the system will provide a detailed answer from its knowledge database.

[0389] An example of a prompt message could be: "Create a sample conversation that provides a travel insurance procedure guide based on the user's emotions. Include examples of guidance tailored to different emotional states."

[0390] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0391] Step 1:

[0392] When a user approaches the device, the device uses its camera to capture the user's face. This facial image data becomes the input, and the device sends this data to the server. The server processes the received image data through a facial recognition algorithm and outputs the user's personal identification information. This process involves data analysis using an image processing library.

[0393] Step 2:

[0394] The server generates a menu of procedures that the user can execute based on the facial recognition results. The input here is the recognized personal identification information. The server refers to the user's attribute information and history to dynamically determine the appropriate procedure menu and outputs that menu to the terminal in JSON format.

[0395] Step 3:

[0396] The terminal displays the procedure menu received from the server on its user interface. The user makes selections using voice or touch, based on this displayed menu. The terminal sends the selected information to the server, triggering the progress of the corresponding procedure.

[0397] Step 4:

[0398] Once the user's procedure selection is transmitted to the server, the server activates the emotion engine and analyzes the user's emotional state. The input for the analysis is the user's facial expressions and voice data. The server passes this data through the emotion recognition system to detect the user's emotional state and outputs the result to the terminal as emotional information.

[0399] Step 5:

[0400] The server adjusts the tone and content of the procedural guidance based on the user's emotional information. Input consists of information necessary for the procedure to progress and emotional data. Based on this information, the server generates reassuring guidance content and outputs it to the terminal.

[0401] Step 6:

[0402] Based on the procedure selected by the user, the server automatically generates the necessary electronic documents from document templates, following prompts. The input is the selected procedure information, and the output is an editable electronic document. This document is sent to the terminal and displayed to the user. The user can edit the electronic document as needed.

[0403] Step 7:

[0404] During the process, if a user asks a question, the audio is converted to text by the device and sent to the server. The server consults a knowledge database, generates an appropriate answer using the transcribed question as input, adjusts the emotional tone, and then outputs the answer to the device.

[0405] Step 8:

[0406] When the process reaches its final stage, the server verifies that all steps have been completed correctly. After verification, the server generates an approval message and sends it to the terminal. The terminal displays this message to the user, informing them that the process has been successfully completed.

[0407] (Application Example 2)

[0408] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0409] In modern self-service systems, information is often provided uniformly without regard for the user's feelings, which frequently leads to confusion and stress. To address this problem, it is necessary to recognize the user's emotional state in real time and provide personalized guidance and answers accordingly, thereby creating a smoother and more effective self-service experience.

[0410] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0411] In this invention, the server includes a function to recognize the user's emotions and adjust the procedure guide and responses accordingly, a function to perform individual authentication based on facial recognition, and a function to accept the user's procedure selection via voice or text. This makes it possible to provide customized procedure guidance and responses according to the user's emotional state.

[0412] "Facial recognition" is a technology that analyzes the facial features of a user to identify them as an individual.

[0413] "Individual authentication" is a procedure used by a system to identify a specific user and grant them access.

[0414] "User" refers to the end user who operates this system.

[0415] "Voice or text" refers to a method of communication used by users to input procedures or questions into a system.

[0416] The "function to accept procedure selection" refers to the ability to provide an interface for users to choose the procedure they wish to perform.

[0417] The "automatic document generation function" is a technology that automatically creates the documents necessary for a procedure based on an algorithm.

[0418] An "FAQ database" is a source of information that compiles frequently asked questions and their answers.

[0419] The "function that recognizes emotions and adjusts procedural guidance and responses accordingly" is a technology that analyzes the user's emotional state and dynamically changes the guidance method and content based on that analysis.

[0420] An "information processing device" is a general term for a series of electronic devices that collect, analyze, and output data.

[0421] The system for realizing this invention is a self-service device that recognizes the user's emotions and adjusts the operation guide and responses accordingly. The system uses a facial recognition camera to capture the user's face when the user approaches the terminal and authenticates a specific individual. This is achieved by utilizing common facial recognition technologies (e.g., OpenCV or the dlib library).

[0422] Subsequently, the user selects the necessary procedure via voice or text input, and the system generates and displays the necessary information for that procedure. This information generation utilizes natural language processing libraries (e.g., spaCy or NLTK). If the user asks a question during the procedure, the voice is converted into text, and the system searches the server's FAQ database for the most appropriate answer and provides it. The FAQ database is often constructed using an SQL database.

[0423] Furthermore, an emotion recognition engine (e.g., Microsoft Azure Emotion API or Google Cloud Vision API) analyzes the user's facial expressions and voice tone, and the guidance method and response tone are adjusted accordingly. This system allows users to use self-services efficiently and stress-free.

[0424] For example, if a customer is using a self-checkout machine in a physical store to purchase a new product and the terminal detects their confused expression, the system will display a helpful message such as, "Is this your first time using this? Do you need assistance?" Additional guidance would also be provided, including detailed instructions on how to use the product and its features.

[0425] An example of a prompt message is, "How can I adjust and display the procedure guide based on the emotions a customer feels when using a self-checkout?"

[0426] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0427] Step 1:

[0428] The device detects the user and captures a facial image using a facial recognition camera. The input is image data acquired by the camera, and the output is the identification information of the recognized user. User authentication is performed using this data, and the individual is identified.

[0429] Step 2:

[0430] The server generates a menu of procedures that the user can select from, based on the authentication result. The input is the user's identification information, and the output is a list of procedure options. The server retrieves the user's history data from the database and presents the most appropriate menu.

[0431] Step 3:

[0432] The user selects their desired procedure on the terminal via voice or text input. The input is the user's voice or text data, and the output is the content of the selected procedure. Voice recognition technology is used to convert the voice into text and analyze the user's selection.

[0433] Step 4:

[0434] The server automatically generates the necessary documents according to the selected procedure and sends them to the terminal. The input is information about the selected procedure, and the output is the generated document data. Natural language processing is used to organize and provide the documents in a way that is easy for the user to understand.

[0435] Step 5:

[0436] The terminal displays documents to the user and provides an editable interface. Input is document data from the server, and output is the user's edited document. The system allows the user to modify the document as needed.

[0437] Step 6:

[0438] When a user asks a question, the audio is converted into text by the device and sent to the server. The input is the user's question in audio form, and the output is the converted question text. An appropriate answer is then selected from the FAQ database.

[0439] Step 7:

[0440] The server analyzes the user's emotional state through an emotion recognition engine, adjusts the tone of its response based on that analysis, and sends the response to the terminal. Input is the user's facial expressions and voice data, and output is the adjusted response. The system communicates flexibly according to the user's emotions.

[0441] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0442] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0443] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0444] [Third Embodiment]

[0445] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0446] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0447] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0448] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0449] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0450] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0451] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0452] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0453] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0454] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0455] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0456] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0457] This invention relates to a system for users to perform self-service procedures using a terminal. The system includes facial recognition for personal authentication, procedure selection using voice and text, automatic document generation, and a question-answering function utilizing an FAQ database. It also features voice and text interfaces to guide the user through the procedure.

[0458] System Operation Overview

[0459] 1. The user faces the device's camera, and facial recognition is performed. This authenticates the user's identity.

[0460] 2. After successful authentication, the device displays a list of available procedures to the user, allowing them to select a procedure via voice or text.

[0461] 3. The server automatically generates the necessary documents based on the procedure selected by the user and sends them to the terminal. A template containing the information to be filled in is used for this purpose.

[0462] 4. The terminal displays documents to the user and provides an interface that allows the user to review or modify the content.

[0463] 5. If the user asks a question during the process, the terminal will convert the question into text and send it to the server.

[0464] 6. The server consults the FAQ database, generates an appropriate answer to the question, and sends it back to the terminal.

[0465] 7. After the document verification is complete, the user selects "Complete Procedure," at which point the server records the completion of the procedure and sends a confirmation notification to the device.

[0466] Specific example

[0467] For example, if a user wants to change the address on their bank account, they first perform facial recognition on the terminal. Then, they select "Change Address" from the menu, and the system automatically generates the necessary documents for the address change. The user enters the required information into the address change form and confirms it. If the user asks "What documents are required for this?" at some point, the system will respond "You will need identification documents and a certificate of residence." Once the documents have been verified and the user has completed the procedure, the system correctly records the procedure and displays a confirmation message to the user. This system allows users to complete the procedure safely and securely 24 hours a day.

[0468] The following describes the processing flow.

[0469] Step 1:

[0470] The user approaches the device and faces the camera for facial recognition. The device uses the camera to capture an image of the user's face.

[0471] Step 2:

[0472] The device sends the acquired facial image to the server and initiates facial recognition. The server uses the facial recognition system to perform personal authentication and returns the result to the device.

[0473] Step 3:

[0474] The device notifies the user that facial recognition was successful and displays a menu of options. The user selects their desired procedure using the device's display or voice input.

[0475] Step 4:

[0476] When the user selects a procedure, the terminal sends that information to the server. The server selects the necessary document templates for the chosen procedure and automatically generates the documents.

[0477] Step 5:

[0478] The server sends the generated document to the terminal. The terminal displays the document to the user and provides a screen where the entered information can be reviewed and modified.

[0479] Step 6:

[0480] If a user asks a question during the process, the terminal converts the voice input into text and sends the question to the server. The server consults the FAQ database, generates an appropriate answer to the question, and returns it to the terminal.

[0481] Step 7:

[0482] The user reviews the documents and selects "Complete Procedure." The device then sends this instruction to the server.

[0483] Step 8:

[0484] The server confirms that the procedure is complete and sends an acknowledgment notification to the terminal. The terminal displays a confirmation message to the user, informing them that the procedure is finished.

[0485] (Example 1)

[0486] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0487] There is a need for self-service systems that allow users to complete procedures at their own pace, without being restricted by time or location. However, conventional systems lack sufficient features for personal authentication and smooth procedures, resulting in a lack of means for users to complete procedures themselves with peace of mind. In particular, there is a lack of mechanisms to respond quickly and appropriately to user questions. Furthermore, delays in responding to questions that arise during the procedure and the burden of complex operations on users are also problems.

[0488] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0489] In this invention, the server includes means for personal authentication based on facial recognition, means for referencing an information database using an AI model, and means for dynamically guiding the user's procedure progress. This enables the user to proceed with the procedure quickly and efficiently. Specifically, by introducing a secure authentication method using facial recognition and providing consistent support from procedure selection to document generation and answering questions, a safe and reliable 24 / 7 self-service environment is realized.

[0490] "Facial recognition" is an authentication method that analyzes facial features to identify an individual and compares them with registered information in a database.

[0491] "Personal authentication" is a process of verifying that a user is who they claim to be, and it is a technology that ensures security by using specific biometric or identification information.

[0492] "Automatic document generation" is a function that automatically inserts necessary information based on a specific template and generates a document in a predetermined format.

[0493] An "information processing device" is a general term for electronic devices used for inputting, processing, and outputting data, and is sometimes commonly referred to as a computer.

[0494] An "information database" is a collection of data in which various types of data are stored in a structured manner, and information can be searched and analyzed through queries.

[0495] An "AI model" is a mathematical framework for performing a specific task using artificial intelligence technology, and its accuracy is improved through learning algorithms.

[0496] A "prompt statement" is an instruction statement that a user enters into the system to perform an intended operation or obtain information, and is used to trigger a specific process.

[0497] One embodiment of the present invention is a system that enables users to perform procedures efficiently and securely through self-service. This system includes facial recognition, procedure selection, automatic document generation, question answering, and progress guidance functions.

[0498] The user first authenticates themselves using facial recognition via the device's camera. The facial recognition technology employs AI-powered facial feature extraction and matching algorithms. This ensures the protection of the user's personal information and provides secure authentication.

[0499] After successful user authentication, the terminal displays a list of available procedures on its menu screen. The user selects their desired procedure using voice input or touch screen operation. This allows the user to easily choose the procedure they need.

[0500] The server automatically generates the necessary documents according to the user's selected procedure. These documents are generated based on pre-configured templates, and by inserting the user's information, the document creation process is expedited.

[0501] Furthermore, the system can utilize a generative AI model to refer to an FAQ database and provide appropriate answers to user questions. In this process, the AI ​​analyzes the question, quickly searches for relevant data, and generates the answer. Natural language processing technology is used in this process.

[0502] Furthermore, throughout the process, the terminal guides the user through voice and text interfaces, helping them to operate safely and intuitively.

[0503] For example, if a user wants to change the address on their bank account, a possible prompt message might be, "Please tell me what documents are required to change the address on my bank account." The system responds to this prompt by providing the necessary documents and procedural details, supporting the user in completing the process smoothly.

[0504] As described above, the present invention enables users to proceed with procedures with peace of mind and efficiency.

[0505] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0506] Step 1:

[0507] The user provides facial recognition input data by facing their face towards the device's camera. The device sends the acquired facial image data to an AI-powered facial recognition engine and performs data processing to match it with personal identification information. If the match is successful, it generates an output indicating that personal authentication is complete and proceeds to the next step.

[0508] Step 2:

[0509] The terminal displays multiple procedure options on the screen to users who have successfully authenticated. This display includes information about available procedures based on the system's database. Users make their selection via voice input or touch operation. The selected procedure information is entered, and an output indicating that the procedure selection is complete is generated.

[0510] Step 3:

[0511] The server selects the necessary document template based on the user's procedure selection, inserts user data, and performs automated document generation. During this process, a template engine is used to process the specified information. The generated document is output to the terminal and provided to the user.

[0512] Step 4:

[0513] The terminal displays the generated document on its screen and provides an interface that allows the user to review and edit the document content. After the user proofreads the document and enters the necessary information, they review the results on the terminal. The terminal generates output indicating that the review is complete and proceeds to the next step.

[0514] Step 5:

[0515] If the user enters a question during the process, the terminal receives the voice or text input and prepares to send the question to the server. This question is treated as a prompt and used as input data to the server.

[0516] Step 6:

[0517] Based on the received prompt, the server utilizes a generative AI model to refer to the FAQ database. This database search performs data calculations to generate the optimal answer to the question. The server then generates output in text format, which is sent to the terminal.

[0518] Step 7:

[0519] When the user chooses to complete the procedure, the server records the completion status and notifies the terminal accordingly. This process inputs a completion notification, and the user receives output that allows them to confirm the completion of the procedure.

[0520] (Application Example 1)

[0521] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0522] In modern online procedures, the complexity of the process is a problem because users are required to take multiple steps. In particular, when registering or updating payment information, the progress of the procedure is often unclear, and it can be difficult to understand what documents and information are needed. There are also challenges in obtaining accurate and prompt answers to user questions. A system that can improve the user experience by solving these problems is needed.

[0523] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0524] In this invention, the server includes means for performing personal authentication based on facial recognition, means for receiving the user's procedure selection by voice or text, means for registering or updating payment information, means for obtaining answers to questions using a natural language generation engine with respect to an information database, and means having a voice and text interface for guiding the procedure progress and payment operations. This enables the automation and efficiency of procedures, allowing users to perform procedures that can be accepted simply and quickly.

[0525] "Facial recognition" is a technology that identifies individuals by analyzing the features of a user's face as digital data.

[0526] "Personal authentication" is a process for verifying that a specific user is who they claim to be.

[0527] "Procedure selection by voice or text" is a feature that allows users to select the procedure they wish to perform through voice input or text input.

[0528] "Methods for automatically generating necessary documents for procedures" refers to technologies that automatically create the necessary procedural documents based on pre-configured templates and information provided by the user.

[0529] A "user interface" is a screen or control panel that provides an efficient and convenient means for a system and a user to exchange information.

[0530] An "information database" is a collection of information used to provide answers to questions.

[0531] "Procedure completion confirmation" is a function that notifies and records to the user that all steps of the selected procedure have been successfully completed.

[0532] "Means for registering or updating payment information" refers to the process of registering a user's new payment information in the system and changing it as needed.

[0533] A "natural language generation engine" is a technology that automatically generates human-readable text in response to input information or questions.

[0534] A "voice and text interface" is a common gateway for users and systems to communicate via voice commands and text messages.

[0535] To implement this invention, a system is needed in which both the server and the user terminal cooperate to enable the user to perform procedures efficiently. First, the server uses facial recognition technology to authenticate the user's identity. This facial recognition process uses an image processing library to acquire facial images in real time using the user terminal's camera. Next, the server uses speech recognition technology to provide an interface for the user to select procedures by voice or text. Here, speech-to-text conversion technology plays a crucial role.

[0536] When a user selects a specific procedure, the server automatically generates the necessary procedural documents using a template engine and sends them to the user's terminal. The user is then provided with an interface on their terminal that allows them to view the documents and edit the information as needed. This interface is designed to be intuitive and easy to use.

[0537] If a user asks a question during the process, the server uses a generative AI model to retrieve the appropriate answer from the information database and provides the user with a response generated in natural language. For example, if a user asks, "What documents are required for this procedure?", the server will respond, "You will need identification documents."

[0538] Furthermore, once the process is complete, the server will inform the user of its completion and provide guidance on how to proceed to the next step. Through this entire process, users can register or update their payment information with confidence.

[0539] An example of a prompt message is, "What information is required when registering a new payment method?", and an accurate answer is generated based on the information database.

[0540] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0541] Step 1:

[0542] The server performs facial recognition of the user. The user faces the camera on their device, and the device captures the facial image and sends it to the server. The server analyzes the received image using an image processing library and compares it with registered facial data to identify the individual. The input is the camera image, and the output is the authentication result.

[0543] Step 2:

[0544] After successful authentication, the server sends a list of available procedures to the terminal. The terminal displays this list and prompts the user to select a procedure via voice or text. The input is the authentication result, and the output is a display of the procedure list.

[0545] Step 3:

[0546] When a user selects a procedure by voice, the terminal converts the voice input into text and sends it to the server. The server analyzes the transmitted text to identify the selected procedure. The input is voice data, and the output is text information of the selected procedure.

[0547] Step 4:

[0548] Based on the specified procedure, the server creates the necessary documents using an automatically generated template. The generated documents are sent to the terminal, which displays them to the user. The user can then review and edit the document contents. The input is the selected procedure information, and the output is the generated document data.

[0549] Step 5:

[0550] If a user asks a question during the process, the terminal transcribes the question into text and sends it to the server. The server uses a generative AI model to search a relevant information database, generates an accurate answer to the question, and sends it back to the terminal. The input is the question text, and the output is the generated answer.

[0551] Step 6:

[0552] When the user selects to complete the procedure, the server verifies that all processes have terminated properly and sends the result to the terminal. The terminal then displays a completion notification to the user. The input is the procedure completion instruction, and the output is the completion notification.

[0553] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0554] This invention is a self-service system incorporating an emotion engine, which aims to automate various procedures and improve convenience for users using a terminal. The emotion engine analyzes the user's facial expressions and voice to recognize their emotional state and utilizes that information for procedural guidance and response generation.

[0555] System Operation Overview

[0556] 1. The user approaches the device and performs facial recognition. The device captures the facial image and sends it to the server for authentication.

[0557] 2. Based on the authentication result, the server generates a menu of procedures that the user can select and sends it to the terminal.

[0558] 3. The terminal displays a procedure menu to the user and provides the ability to select a procedure by voice or text. At this point, the emotion engine operates, analyzing the user's facial expressions and tone of voice to recognize their emotions.

[0559] 4. Recognized emotional information is reflected in the content and tone of the procedural guide. If the user is experiencing stress, the guide will be adjusted to provide a more gentle approach.

[0560] 5. Based on the user's selection, the server automatically generates the necessary documents and sends them to the terminal. The terminal displays the documents to the user and provides an editable interface.

[0561] 6. When a question arises, the user's voice is transcribed into text and sent to the server. An appropriate answer is generated by referring to the FAQ database. Here again, sentiment information is used to adjust the tone and detail of the answer.

[0562] 7. Once the procedure is complete, the server will perform a final check and send and display an approval message to the terminal.

[0563] Specific example

[0564] For example, when a user is going through the process of purchasing travel insurance, if the terminal detects an anxious expression on their face, the system will guide them through the process in a calming tone. It will then generate a helpful inquiry such as, "Is there anything I can help you with?" to support the user. If the user asks, "Does this insurance cover pre-existing conditions?", the system will provide a detailed answer from the FAQ, adjusting the tone of the explanation according to the user's emotions. In this way, a user-friendly self-service can be provided.

[0565] The following describes the processing flow.

[0566] Step 1:

[0567] The user faces the device, initiating facial recognition. The device captures an image of the user's face using its camera and sends the image data to the server for personal authentication.

[0568] Step 2:

[0569] The server processes the received facial image data using its facial recognition system and sends the authentication result back to the terminal. If authentication is successful, the process proceeds to the next step.

[0570] Step 3:

[0571] The terminal notifies the user of successful authentication and displays a procedure menu. The user selects their desired procedure using voice or the touch panel. The emotion engine analyzes the user's facial expressions and voice to determine their emotions.

[0572] Step 4:

[0573] The server identifies the necessary document templates based on the user's selected procedure and automatically generates and sends the documents to the terminal. The content and tone of the procedure guidance are adjusted based on the recognized emotions.

[0574] Step 5:

[0575] The terminal provides the user with an interface for viewing, editing, and reviewing documents. If the user experiences anxiety or questions during the process, the emotion engine detects this and prepares to provide appropriate guidance.

[0576] Step 6:

[0577] When a user asks a question, the device converts the voice input into text and sends it to the server. The server searches the FAQ database and generates the most relevant answer. The tone and detail of the answer are adjusted according to the user's mood.

[0578] Step 7:

[0579] Once the user completes the procedure, the device sends that information to the server. The server confirms that the procedure was completed successfully and sends a final confirmation message to the device to notify the user.

[0580] Step 8:

[0581] The device displays a final confirmation message to inform the user that the procedure is complete. After confirmation, the user can leave the device.

[0582] (Example 2)

[0583] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0584] Current self-service systems lack consideration for users' emotional states, making it difficult for users to perform procedures with peace of mind. Furthermore, the lack of dynamic adjustments to the procedure flow based on individual user circumstances leads to decreased operational efficiency and user satisfaction.

[0585] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0586] In this invention, the server includes means for identifying an individual based on facial recognition data, means for accepting process selections based on the user's speech or input information, and means for analyzing the user's emotional state and adjusting the guidance content. This enables responses that take the user's emotions into consideration and provides a flexible procedural flow tailored to individual situations.

[0587] "Facial recognition data" refers to information extracted from facial images taken to identify individual characteristics.

[0588] "Means of identifying individuals" refers to processes and technologies that use facial recognition data to distinguish a specific individual from others.

[0589] "User utterances or input information" refers to voice or text data used by users to communicate their choices or intentions to the system.

[0590] A "means for accepting process selection" refers to a mechanism for receiving input related to the user's procedure selection and initiating processing based on that input.

[0591] An "electronic document" is a digital document containing information necessary for a process, which can be created, edited, and displayed on a computer.

[0592] A "knowledge database" is a collection of data that stores information and answers in a specific field, and is used to respond appropriately to questions from users.

[0593] "Means of analyzing emotional states" refers to technical methods that analyze a user's facial expressions and voice to determine their emotions.

[0594] "Means for adjusting guidance content" refers to a mechanism for changing the content and tone of guidance provided to the user based on their analyzed emotional state.

[0595] This invention provides a self-service system incorporating an emotion engine, allowing users to automate various procedures and improve convenience through their user terminals. The system performs personal identification based on user facial recognition data and analyzes their emotional state to provide flexible guidance tailored to the user's situation.

[0596] Specifically, when a user approaches the terminal, the facial recognition system activates and captures an image of the user's face. The hardware used includes a high-performance camera, and the software includes image processing libraries such as OpenCV. The terminal sends the captured facial image data to the server for personal identification processing. Machine learning algorithms are used in this identification process. If facial recognition is successful, the server creates a procedural menu related to the user and sends it to the terminal.

[0597] Users can select their desired task from a procedure menu displayed on the terminal using voice or text input. In this step, an emotion engine analyzes the user's facial expressions and voice tone to evaluate the user's emotional state. This analysis uses a Python-based library, such as the Emotion Recognition API. The tone and content of the guidance are adjusted according to the evaluated emotional state. For example, if the user is feeling anxious, the terminal will provide a friendly voice guidance to calm the user.

[0598] Once a procedure is selected, the server automatically generates the necessary electronic documents using document templates. These generated electronic documents are displayed on the terminal in an editable format. If the user asks additional questions, these questions are transcribed using speech recognition technology, and answers are generated by referencing the server's knowledge database. This ensures that an appropriate response is selected, taking the user's feelings into consideration.

[0599] For example, when a user is processing travel insurance on a terminal, if the emotion engine detects the user's anxiety, it will add calming phrases such as "Is there anything I can help you with?" in its guidance. Next, if the user asks a question, either by typing or speaking, such as "Please tell me what is covered by this insurance," the system will provide a detailed answer from its knowledge database.

[0600] An example of a prompt message could be: "Create a sample conversation that provides a travel insurance procedure guide based on the user's emotions. Include examples of guidance tailored to different emotional states."

[0601] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0602] Step 1:

[0603] When a user approaches the device, the device uses its camera to capture the user's face. This facial image data becomes the input, and the device sends this data to the server. The server processes the received image data through a facial recognition algorithm and outputs the user's personal identification information. This process involves data analysis using an image processing library.

[0604] Step 2:

[0605] The server generates a menu of procedures that the user can execute based on the facial recognition results. The input here is the recognized personal identification information. The server refers to the user's attribute information and history to dynamically determine the appropriate procedure menu and outputs that menu to the terminal in JSON format.

[0606] Step 3:

[0607] The terminal displays the procedure menu received from the server on its user interface. The user makes selections using voice or touch, based on this displayed menu. The terminal sends the selected information to the server, triggering the progress of the corresponding procedure.

[0608] Step 4:

[0609] Once the user's procedure selection is transmitted to the server, the server activates the emotion engine and analyzes the user's emotional state. The input for the analysis is the user's facial expressions and voice data. The server passes this data through the emotion recognition system to detect the user's emotional state and outputs the result to the terminal as emotional information.

[0610] Step 5:

[0611] The server adjusts the tone and content of the procedural guidance based on the user's emotional information. Input consists of information necessary for the procedure to progress and emotional data. Based on this information, the server generates reassuring guidance content and outputs it to the terminal.

[0612] Step 6:

[0613] Based on the procedure selected by the user, the server automatically generates the necessary electronic documents from document templates, following prompts. The input is the selected procedure information, and the output is an editable electronic document. This document is sent to the terminal and displayed to the user. The user can edit the electronic document as needed.

[0614] Step 7:

[0615] During the process, if a user asks a question, the audio is converted to text by the device and sent to the server. The server consults a knowledge database, generates an appropriate answer using the transcribed question as input, adjusts the emotional tone, and then outputs the answer to the device.

[0616] Step 8:

[0617] When the process reaches its final stage, the server verifies that all steps have been completed correctly. After verification, the server generates an approval message and sends it to the terminal. The terminal displays this message to the user, informing them that the process has been successfully completed.

[0618] (Application Example 2)

[0619] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0620] In modern self-service systems, information is often provided uniformly without regard for the user's feelings, which frequently leads to confusion and stress. To address this problem, it is necessary to recognize the user's emotional state in real time and provide personalized guidance and answers accordingly, thereby creating a smoother and more effective self-service experience.

[0621] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0622] In this invention, the server includes a function to recognize the user's emotions and adjust the procedure guide and responses accordingly, a function to perform individual authentication based on facial recognition, and a function to accept the user's procedure selection via voice or text. This makes it possible to provide customized procedure guidance and responses according to the user's emotional state.

[0623] "Facial recognition" is a technology that analyzes the facial features of a user to identify them as an individual.

[0624] "Individual authentication" is a procedure used by a system to identify a specific user and grant them access.

[0625] "User" refers to the end user who operates this system.

[0626] "Voice or text" refers to a method of communication used by users to input procedures or questions into a system.

[0627] The "function to accept procedure selection" refers to the ability to provide an interface for users to choose the procedure they wish to perform.

[0628] The "automatic document generation function" is a technology that automatically creates the documents necessary for a procedure based on an algorithm.

[0629] An "FAQ database" is a source of information that compiles frequently asked questions and their answers.

[0630] The "function that recognizes emotions and adjusts procedural guidance and responses accordingly" is a technology that analyzes the user's emotional state and dynamically changes the guidance method and content based on that analysis.

[0631] An "information processing device" is a general term for a series of electronic devices that collect, analyze, and output data.

[0632] The system for realizing this invention is a self-service device that recognizes the user's emotions and adjusts the operation guide and responses accordingly. The system uses a facial recognition camera to capture the user's face when the user approaches the terminal and authenticates a specific individual. This is achieved by utilizing common facial recognition technologies (e.g., OpenCV or the dlib library).

[0633] Subsequently, the user selects the necessary procedure via voice or text input, and the system generates and displays the necessary information for that procedure. This information generation utilizes natural language processing libraries (e.g., spaCy or NLTK). If the user asks a question during the procedure, the voice is converted into text, and the system searches the server's FAQ database for the most appropriate answer and provides it. The FAQ database is often constructed using an SQL database.

[0634] Furthermore, an emotion recognition engine (e.g., Microsoft Azure Emotion API or Google Cloud Vision API) analyzes the user's facial expressions and voice tone, and the guidance method and response tone are adjusted accordingly. This system allows users to use self-services efficiently and stress-free.

[0635] For example, if a customer is using a self-checkout machine in a physical store to purchase a new product and the terminal detects their confused expression, the system will display a helpful message such as, "Is this your first time using this? Do you need assistance?" Additional guidance would also be provided, including detailed instructions on how to use the product and its features.

[0636] An example of a prompt message is, "How can I adjust and display the procedure guide based on the emotions a customer feels when using a self-checkout?"

[0637] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0638] Step 1:

[0639] The device detects the user and captures a facial image using a facial recognition camera. The input is image data acquired by the camera, and the output is the identification information of the recognized user. User authentication is performed using this data, and the individual is identified.

[0640] Step 2:

[0641] The server generates a menu of procedures that the user can select from, based on the authentication result. The input is the user's identification information, and the output is a list of procedure options. The server retrieves the user's history data from the database and presents the most appropriate menu.

[0642] Step 3:

[0643] The user selects their desired procedure on the terminal via voice or text input. The input is the user's voice or text data, and the output is the content of the selected procedure. Voice recognition technology is used to convert the voice into text and analyze the user's selection.

[0644] Step 4:

[0645] The server automatically generates the necessary documents according to the selected procedure and sends them to the terminal. The input is information about the selected procedure, and the output is the generated document data. Natural language processing is used to organize and provide the documents in a way that is easy for the user to understand.

[0646] Step 5:

[0647] The terminal displays documents to the user and provides an editable interface. Input is document data from the server, and output is the user's edited document. The system allows the user to modify the document as needed.

[0648] Step 6:

[0649] When a user asks a question, the audio is converted into text by the device and sent to the server. The input is the user's question in audio form, and the output is the converted question text. An appropriate answer is then selected from the FAQ database.

[0650] Step 7:

[0651] The server analyzes the user's emotional state through an emotion recognition engine, adjusts the tone of its response based on that analysis, and sends the response to the terminal. Input is the user's facial expressions and voice data, and output is the adjusted response. The system communicates flexibly according to the user's emotions.

[0652] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0653] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0654] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0655] [Fourth Embodiment]

[0656] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0657] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0658] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0659] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0660] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0661] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0662] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0663] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0664] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0665] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0666] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0667] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0668] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0669] This invention relates to a system for users to perform self-service procedures using a terminal. The system includes facial recognition for personal authentication, procedure selection using voice and text, automatic document generation, and a question-answering function utilizing an FAQ database. It also features voice and text interfaces to guide the user through the procedure.

[0670] System Operation Overview

[0671] 1. The user faces the device's camera, and facial recognition is performed. This authenticates the user's identity.

[0672] 2. After successful authentication, the device displays a list of available procedures to the user, allowing them to select a procedure via voice or text.

[0673] 3. The server automatically generates the necessary documents based on the procedure selected by the user and sends them to the terminal. A template containing the information to be filled in is used for this purpose.

[0674] 4. The terminal displays documents to the user and provides an interface that allows the user to review or modify the content.

[0675] 5. If the user asks a question during the process, the terminal will convert the question into text and send it to the server.

[0676] 6. The server consults the FAQ database, generates an appropriate answer to the question, and sends it back to the terminal.

[0677] 7. After the document verification is complete, the user selects "Complete Procedure," at which point the server records the completion of the procedure and sends a confirmation notification to the device.

[0678] Specific example

[0679] For example, if a user wants to change the address on their bank account, they first perform facial recognition on the terminal. Then, they select "Change Address" from the menu, and the system automatically generates the necessary documents for the address change. The user enters the required information into the address change form and confirms it. If the user asks "What documents are required for this?" at some point, the system will respond "You will need identification documents and a certificate of residence." Once the documents have been verified and the user has completed the procedure, the system correctly records the procedure and displays a confirmation message to the user. This system allows users to complete the procedure safely and securely 24 hours a day.

[0680] The following describes the processing flow.

[0681] Step 1:

[0682] The user approaches the device and faces the camera for facial recognition. The device uses the camera to capture an image of the user's face.

[0683] Step 2:

[0684] The device sends the acquired facial image to the server and initiates facial recognition. The server uses the facial recognition system to perform personal authentication and returns the result to the device.

[0685] Step 3:

[0686] The device notifies the user that facial recognition was successful and displays a menu of options. The user selects their desired procedure using the device's display or voice input.

[0687] Step 4:

[0688] When the user selects a procedure, the terminal sends that information to the server. The server selects the necessary document templates for the chosen procedure and automatically generates the documents.

[0689] Step 5:

[0690] The server sends the generated document to the terminal. The terminal displays the document to the user and provides a screen where the entered information can be reviewed and modified.

[0691] Step 6:

[0692] If a user asks a question during the process, the terminal converts the voice input into text and sends the question to the server. The server consults the FAQ database, generates an appropriate answer to the question, and returns it to the terminal.

[0693] Step 7:

[0694] The user reviews the documents and selects "Complete Procedure." The device then sends this instruction to the server.

[0695] Step 8:

[0696] The server confirms that the procedure is complete and sends an acknowledgment notification to the terminal. The terminal displays a confirmation message to the user, informing them that the procedure is finished.

[0697] (Example 1)

[0698] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0699] There is a need for self-service systems that allow users to complete procedures at their own pace, without being restricted by time or location. However, conventional systems lack sufficient features for personal authentication and smooth procedures, resulting in a lack of means for users to complete procedures themselves with peace of mind. In particular, there is a lack of mechanisms to respond quickly and appropriately to user questions. Furthermore, delays in responding to questions that arise during the procedure and the burden of complex operations on users are also problems.

[0700] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0701] In this invention, the server includes means for personal authentication based on facial recognition, means for referencing an information database using an AI model, and means for dynamically guiding the user's procedure progress. This enables the user to proceed with the procedure quickly and efficiently. Specifically, by introducing a secure authentication method using facial recognition and providing consistent support from procedure selection to document generation and answering questions, a safe and reliable 24 / 7 self-service environment is realized.

[0702] "Facial recognition" is an authentication method that analyzes facial features to identify an individual and compares them with registered information in a database.

[0703] "Personal authentication" is a process of verifying that a user is who they claim to be, and it is a technology that ensures security by using specific biometric or identification information.

[0704] "Automatic document generation" is a function that automatically inserts necessary information based on a specific template and generates a document in a predetermined format.

[0705] An "information processing device" is a general term for electronic devices used for inputting, processing, and outputting data, and is sometimes commonly referred to as a computer.

[0706] An "information database" is a collection of data in which various types of data are stored in a structured manner, and information can be searched and analyzed through queries.

[0707] An "AI model" is a mathematical framework for performing a specific task using artificial intelligence technology, and its accuracy is improved through learning algorithms.

[0708] A "prompt statement" is an instruction statement that a user enters into the system to perform an intended operation or obtain information, and is used to trigger a specific process.

[0709] One embodiment of the present invention is a system that enables users to perform procedures efficiently and securely through self-service. This system includes facial recognition, procedure selection, automatic document generation, question answering, and progress guidance functions.

[0710] The user first authenticates themselves using facial recognition via the device's camera. The facial recognition technology employs AI-powered facial feature extraction and matching algorithms. This ensures the protection of the user's personal information and provides secure authentication.

[0711] After successful user authentication, the terminal displays a list of available procedures on its menu screen. The user selects their desired procedure using voice input or touch screen operation. This allows the user to easily choose the procedure they need.

[0712] The server automatically generates the necessary documents according to the user's selected procedure. These documents are generated based on pre-configured templates, and by inserting the user's information, the document creation process is expedited.

[0713] Furthermore, the system can utilize a generative AI model to refer to an FAQ database and provide appropriate answers to user questions. In this process, the AI ​​analyzes the question, quickly searches for relevant data, and generates the answer. Natural language processing technology is used in this process.

[0714] Furthermore, throughout the process, the terminal guides the user through voice and text interfaces, helping them to operate safely and intuitively.

[0715] For example, if a user wants to change the address on their bank account, a possible prompt message might be, "Please tell me what documents are required to change the address on my bank account." The system responds to this prompt by providing the necessary documents and procedural details, supporting the user in completing the process smoothly.

[0716] As described above, the present invention enables users to proceed with procedures with peace of mind and efficiency.

[0717] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0718] Step 1:

[0719] The user provides facial recognition input data by facing their face towards the device's camera. The device sends the acquired facial image data to an AI-powered facial recognition engine and performs data processing to match it with personal identification information. If the match is successful, it generates an output indicating that personal authentication is complete and proceeds to the next step.

[0720] Step 2:

[0721] The terminal displays multiple procedure options on the screen to users who have successfully authenticated. This display includes information about available procedures based on the system's database. Users make their selection via voice input or touch operation. The selected procedure information is entered, and an output indicating that the procedure selection is complete is generated.

[0722] Step 3:

[0723] The server selects the necessary document template based on the user's procedure selection, inserts user data, and performs automated document generation. During this process, a template engine is used to process the specified information. The generated document is output to the terminal and provided to the user.

[0724] Step 4:

[0725] The terminal displays the generated document on its screen and provides an interface that allows the user to review and edit the document content. After the user proofreads the document and enters the necessary information, they review the results on the terminal. The terminal generates output indicating that the review is complete and proceeds to the next step.

[0726] Step 5:

[0727] If the user enters a question during the process, the terminal receives the voice or text input and prepares to send the question to the server. This question is treated as a prompt and used as input data to the server.

[0728] Step 6:

[0729] Based on the received prompt, the server utilizes a generative AI model to refer to the FAQ database. This database search performs data calculations to generate the optimal answer to the question. The server then generates output in text format, which is sent to the terminal.

[0730] Step 7:

[0731] When the user chooses to complete the procedure, the server records the completion status and notifies the terminal accordingly. This process inputs a completion notification, and the user receives output that allows them to confirm the completion of the procedure.

[0732] (Application Example 1)

[0733] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0734] In modern online procedures, the complexity of the process is a problem because users are required to take multiple steps. In particular, when registering or updating payment information, the progress of the procedure is often unclear, and it can be difficult to understand what documents and information are needed. There are also challenges in obtaining accurate and prompt answers to user questions. A system that can improve the user experience by solving these problems is needed.

[0735] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0736] In this invention, the server includes means for performing personal authentication based on facial recognition, means for receiving the user's procedure selection by voice or text, means for registering or updating payment information, means for obtaining answers to questions using a natural language generation engine with respect to an information database, and means having a voice and text interface for guiding the procedure progress and payment operations. This enables the automation and efficiency of procedures, allowing users to perform procedures that can be accepted simply and quickly.

[0737] "Facial recognition" is a technology that identifies individuals by analyzing the features of a user's face as digital data.

[0738] "Personal authentication" is a process for verifying that a specific user is who they claim to be.

[0739] "Procedure selection by voice or text" is a feature that allows users to select the procedure they wish to perform through voice input or text input.

[0740] "Methods for automatically generating necessary documents for procedures" refers to technologies that automatically create the necessary procedural documents based on pre-configured templates and information provided by the user.

[0741] A "user interface" is a screen or control panel that provides an efficient and convenient means for a system and a user to exchange information.

[0742] An "information database" is a collection of information used to provide answers to questions.

[0743] "Procedure completion confirmation" is a function that notifies and records to the user that all steps of the selected procedure have been successfully completed.

[0744] "Means for registering or updating payment information" refers to the process of registering a user's new payment information in the system and changing it as needed.

[0745] A "natural language generation engine" is a technology that automatically generates human-readable text in response to input information or questions.

[0746] A "voice and text interface" is a common gateway for users and systems to communicate via voice commands and text messages.

[0747] To implement this invention, a system is needed in which both the server and the user terminal cooperate to enable the user to perform procedures efficiently. First, the server uses facial recognition technology to authenticate the user's identity. This facial recognition process uses an image processing library to acquire facial images in real time using the user terminal's camera. Next, the server uses speech recognition technology to provide an interface for the user to select procedures by voice or text. Here, speech-to-text conversion technology plays a crucial role.

[0748] When a user selects a specific procedure, the server automatically generates the necessary procedural documents using a template engine and sends them to the user's terminal. The user is then provided with an interface on their terminal that allows them to view the documents and edit the information as needed. This interface is designed to be intuitive and easy to use.

[0749] If a user asks a question during the process, the server uses a generative AI model to retrieve the appropriate answer from the information database and provides the user with a response generated in natural language. For example, if a user asks, "What documents are required for this procedure?", the server will respond, "You will need identification documents."

[0750] Furthermore, once the process is complete, the server will inform the user of its completion and provide guidance on how to proceed to the next step. Through this entire process, users can register or update their payment information with confidence.

[0751] An example of a prompt message is, "What information is required when registering a new payment method?", and an accurate answer is generated based on the information database.

[0752] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0753] Step 1:

[0754] The server performs facial recognition of the user. The user faces the camera on their device, and the device captures the facial image and sends it to the server. The server analyzes the received image using an image processing library and compares it with registered facial data to identify the individual. The input is the camera image, and the output is the authentication result.

[0755] Step 2:

[0756] After successful authentication, the server sends a list of available procedures to the terminal. The terminal displays this list and prompts the user to select a procedure via voice or text. The input is the authentication result, and the output is a display of the procedure list.

[0757] Step 3:

[0758] When a user selects a procedure by voice, the terminal converts the voice input into text and sends it to the server. The server analyzes the transmitted text to identify the selected procedure. The input is voice data, and the output is text information of the selected procedure.

[0759] Step 4:

[0760] Based on the specified procedure, the server creates the necessary documents using an automatically generated template. The generated documents are sent to the terminal, which displays them to the user. The user can then review and edit the document contents. The input is the selected procedure information, and the output is the generated document data.

[0761] Step 5:

[0762] If a user asks a question during the process, the terminal transcribes the question into text and sends it to the server. The server uses a generative AI model to search a relevant information database, generates an accurate answer to the question, and sends it back to the terminal. The input is the question text, and the output is the generated answer.

[0763] Step 6:

[0764] When the user selects to complete the procedure, the server verifies that all processes have terminated properly and sends the result to the terminal. The terminal then displays a completion notification to the user. The input is the procedure completion instruction, and the output is the completion notification.

[0765] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0766] This invention is a self-service system incorporating an emotion engine, which aims to automate various procedures and improve convenience for users using a terminal. The emotion engine analyzes the user's facial expressions and voice to recognize their emotional state and utilizes that information for procedural guidance and response generation.

[0767] System Operation Overview

[0768] 1. The user approaches the device and performs facial recognition. The device captures the facial image and sends it to the server for authentication.

[0769] 2. Based on the authentication result, the server generates a menu of procedures that the user can select and sends it to the terminal.

[0770] 3. The terminal displays a procedure menu to the user and provides the ability to select a procedure by voice or text. At this point, the emotion engine operates, analyzing the user's facial expressions and tone of voice to recognize their emotions.

[0771] 4. Recognized emotional information is reflected in the content and tone of the procedural guide. If the user is experiencing stress, the guide will be adjusted to provide a more gentle approach.

[0772] 5. Based on the user's selection, the server automatically generates the necessary documents and sends them to the terminal. The terminal displays the documents to the user and provides an editable interface.

[0773] 6. When a question arises, the user's voice is transcribed into text and sent to the server. An appropriate answer is generated by referring to the FAQ database. Here again, sentiment information is used to adjust the tone and detail of the answer.

[0774] 7. Once the procedure is complete, the server will perform a final check and send and display an approval message to the terminal.

[0775] Specific example

[0776] For example, when a user is going through the process of purchasing travel insurance, if the terminal detects an anxious expression on their face, the system will guide them through the process in a calming tone. It will then generate a helpful inquiry such as, "Is there anything I can help you with?" to support the user. If the user asks, "Does this insurance cover pre-existing conditions?", the system will provide a detailed answer from the FAQ, adjusting the tone of the explanation according to the user's emotions. In this way, a user-friendly self-service can be provided.

[0777] The following describes the processing flow.

[0778] Step 1:

[0779] The user faces the device, initiating facial recognition. The device captures an image of the user's face using its camera and sends the image data to the server for personal authentication.

[0780] Step 2:

[0781] The server processes the received facial image data using its facial recognition system and sends the authentication result back to the terminal. If authentication is successful, the process proceeds to the next step.

[0782] Step 3:

[0783] The terminal notifies the user of successful authentication and displays a procedure menu. The user selects their desired procedure using voice or the touch panel. The emotion engine analyzes the user's facial expressions and voice to determine their emotions.

[0784] Step 4:

[0785] The server identifies the necessary document templates based on the user's selected procedure and automatically generates and sends the documents to the terminal. The content and tone of the procedure guidance are adjusted based on the recognized emotions.

[0786] Step 5:

[0787] The terminal provides the user with an interface for viewing, editing, and reviewing documents. If the user experiences anxiety or questions during the process, the emotion engine detects this and prepares to provide appropriate guidance.

[0788] Step 6:

[0789] When a user asks a question, the device converts the voice input into text and sends it to the server. The server searches the FAQ database and generates the most relevant answer. The tone and detail of the answer are adjusted according to the user's mood.

[0790] Step 7:

[0791] Once the user completes the procedure, the device sends that information to the server. The server confirms that the procedure was completed successfully and sends a final confirmation message to the device to notify the user.

[0792] Step 8:

[0793] The device displays a final confirmation message to inform the user that the procedure is complete. After confirmation, the user can leave the device.

[0794] (Example 2)

[0795] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0796] Current self-service systems lack consideration for users' emotional states, making it difficult for users to perform procedures with peace of mind. Furthermore, the lack of dynamic adjustments to the procedure flow based on individual user circumstances leads to decreased operational efficiency and user satisfaction.

[0797] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0798] In this invention, the server includes means for identifying an individual based on facial recognition data, means for accepting process selections based on the user's speech or input information, and means for analyzing the user's emotional state and adjusting the guidance content. This enables responses that take the user's emotions into consideration and provides a flexible procedural flow tailored to individual situations.

[0799] "Facial recognition data" refers to information extracted from facial images taken to identify individual characteristics.

[0800] "Means of identifying individuals" refers to processes and technologies that use facial recognition data to distinguish a specific individual from others.

[0801] "User utterances or input information" refers to voice or text data used by users to communicate their choices or intentions to the system.

[0802] A "means for accepting process selection" refers to a mechanism for receiving input related to the user's procedure selection and initiating processing based on that input.

[0803] An "electronic document" is a digital document containing information necessary for a process, which can be created, edited, and displayed on a computer.

[0804] A "knowledge database" is a collection of data that stores information and answers in a specific field, and is used to respond appropriately to questions from users.

[0805] "Means of analyzing emotional states" refers to technical methods that analyze a user's facial expressions and voice to determine their emotions.

[0806] "Means for adjusting guidance content" refers to a mechanism for changing the content and tone of guidance provided to the user based on their analyzed emotional state.

[0807] This invention provides a self-service system incorporating an emotion engine, allowing users to automate various procedures and improve convenience through their user terminals. The system performs personal identification based on user facial recognition data and analyzes their emotional state to provide flexible guidance tailored to the user's situation.

[0808] Specifically, when a user approaches the terminal, the facial recognition system activates and captures an image of the user's face. The hardware used includes a high-performance camera, and the software includes image processing libraries such as OpenCV. The terminal sends the captured facial image data to the server for personal identification processing. Machine learning algorithms are used in this identification process. If facial recognition is successful, the server creates a procedural menu related to the user and sends it to the terminal.

[0809] Users can select their desired task from a procedure menu displayed on the terminal using voice or text input. In this step, an emotion engine analyzes the user's facial expressions and voice tone to evaluate the user's emotional state. This analysis uses a Python-based library, such as the Emotion Recognition API. The tone and content of the guidance are adjusted according to the evaluated emotional state. For example, if the user is feeling anxious, the terminal will provide a friendly voice guidance to calm the user.

[0810] Once a procedure is selected, the server automatically generates the necessary electronic documents using document templates. These generated electronic documents are displayed on the terminal in an editable format. If the user asks additional questions, these questions are transcribed using speech recognition technology, and answers are generated by referencing the server's knowledge database. This ensures that an appropriate response is selected, taking the user's feelings into consideration.

[0811] For example, when a user is processing travel insurance on a terminal, if the emotion engine detects the user's anxiety, it will add calming phrases such as "Is there anything I can help you with?" in its guidance. Next, if the user asks a question, either by typing or speaking, such as "Please tell me what is covered by this insurance," the system will provide a detailed answer from its knowledge database.

[0812] An example of a prompt message could be: "Create a sample conversation that provides a travel insurance procedure guide based on the user's emotions. Include examples of guidance tailored to different emotional states."

[0813] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0814] Step 1:

[0815] When a user approaches the device, the device uses its camera to capture the user's face. This facial image data becomes the input, and the device sends this data to the server. The server processes the received image data through a facial recognition algorithm and outputs the user's personal identification information. This process involves data analysis using an image processing library.

[0816] Step 2:

[0817] The server generates a menu of procedures that the user can execute based on the facial recognition results. The input here is the recognized personal identification information. The server refers to the user's attribute information and history to dynamically determine the appropriate procedure menu and outputs that menu to the terminal in JSON format.

[0818] Step 3:

[0819] The terminal displays the procedure menu received from the server on its user interface. The user makes selections using voice or touch, based on this displayed menu. The terminal sends the selected information to the server, triggering the progress of the corresponding procedure.

[0820] Step 4:

[0821] Once the user's procedure selection is transmitted to the server, the server activates the emotion engine and analyzes the user's emotional state. The input for the analysis is the user's facial expressions and voice data. The server passes this data through the emotion recognition system to detect the user's emotional state and outputs the result to the terminal as emotional information.

[0822] Step 5:

[0823] The server adjusts the tone and content of the procedural guidance based on the user's emotional information. Input consists of information necessary for the procedure to progress and emotional data. Based on this information, the server generates reassuring guidance content and outputs it to the terminal.

[0824] Step 6:

[0825] Based on the procedure selected by the user, the server automatically generates the necessary electronic documents from document templates, following prompts. The input is the selected procedure information, and the output is an editable electronic document. This document is sent to the terminal and displayed to the user. The user can edit the electronic document as needed.

[0826] Step 7:

[0827] During the process, if a user asks a question, the audio is converted to text by the device and sent to the server. The server consults a knowledge database, generates an appropriate answer using the transcribed question as input, adjusts the emotional tone, and then outputs the answer to the device.

[0828] Step 8:

[0829] When the process reaches its final stage, the server verifies that all steps have been completed correctly. After verification, the server generates an approval message and sends it to the terminal. The terminal displays this message to the user, informing them that the process has been successfully completed.

[0830] (Application Example 2)

[0831] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0832] In modern self-service systems, information is often provided uniformly without regard for the user's feelings, which frequently leads to confusion and stress. To address this problem, it is necessary to recognize the user's emotional state in real time and provide personalized guidance and answers accordingly, thereby creating a smoother and more effective self-service experience.

[0833] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0834] In this invention, the server includes a function to recognize the user's emotions and adjust the procedure guide and responses accordingly, a function to perform individual authentication based on facial recognition, and a function to accept the user's procedure selection via voice or text. This makes it possible to provide customized procedure guidance and responses according to the user's emotional state.

[0835] "Facial recognition" is a technology that analyzes the facial features of a user to identify them as an individual.

[0836] "Individual authentication" is a procedure used by a system to identify a specific user and grant them access.

[0837] "User" refers to the end user who operates this system.

[0838] "Voice or text" refers to a method of communication used by users to input procedures or questions into a system.

[0839] The "function to accept procedure selection" refers to the ability to provide an interface for users to choose the procedure they wish to perform.

[0840] The "automatic document generation function" is a technology that automatically creates the documents necessary for a procedure based on an algorithm.

[0841] An "FAQ database" is a source of information that compiles frequently asked questions and their answers.

[0842] The "function that recognizes emotions and adjusts procedural guidance and responses accordingly" is a technology that analyzes the user's emotional state and dynamically changes the guidance method and content based on that analysis.

[0843] An "information processing device" is a general term for a series of electronic devices that collect, analyze, and output data.

[0844] The system for realizing this invention is a self-service device that recognizes the user's emotions and adjusts the operation guide and responses accordingly. The system uses a facial recognition camera to capture the user's face when the user approaches the terminal and authenticates a specific individual. This is achieved by utilizing common facial recognition technologies (e.g., OpenCV or the dlib library).

[0845] Subsequently, the user selects the necessary procedure via voice or text input, and the system generates and displays the necessary information for that procedure. This information generation utilizes natural language processing libraries (e.g., spaCy or NLTK). If the user asks a question during the procedure, the voice is converted into text, and the system searches the server's FAQ database for the most appropriate answer and provides it. The FAQ database is often constructed using an SQL database.

[0846] Furthermore, an emotion recognition engine (e.g., Microsoft Azure Emotion API or Google Cloud Vision API) analyzes the user's facial expressions and voice tone, and the guidance method and response tone are adjusted accordingly. This system allows users to use self-services efficiently and stress-free.

[0847] For example, if a customer is using a self-checkout machine in a physical store to purchase a new product and the terminal detects their confused expression, the system will display a helpful message such as, "Is this your first time using this? Do you need assistance?" Additional guidance would also be provided, including detailed instructions on how to use the product and its features.

[0848] An example of a prompt message is, "How can I adjust and display the procedure guide based on the emotions a customer feels when using a self-checkout?"

[0849] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0850] Step 1:

[0851] The device detects the user and captures a facial image using a facial recognition camera. The input is image data acquired by the camera, and the output is the identification information of the recognized user. User authentication is performed using this data, and the individual is identified.

[0852] Step 2:

[0853] The server generates a menu of procedures that the user can select from, based on the authentication result. The input is the user's identification information, and the output is a list of procedure options. The server retrieves the user's history data from the database and presents the most appropriate menu.

[0854] Step 3:

[0855] The user selects their desired procedure on the terminal via voice or text input. The input is the user's voice or text data, and the output is the content of the selected procedure. Voice recognition technology is used to convert the voice into text and analyze the user's selection.

[0856] Step 4:

[0857] The server automatically generates the necessary documents according to the selected procedure and sends them to the terminal. The input is information about the selected procedure, and the output is the generated document data. Natural language processing is used to organize and provide the documents in a way that is easy for the user to understand.

[0858] Step 5:

[0859] The terminal displays documents to the user and provides an editable interface. Input is document data from the server, and output is the user's edited document. The system allows the user to modify the document as needed.

[0860] Step 6:

[0861] When a user asks a question, the audio is converted into text by the device and sent to the server. The input is the user's question in audio form, and the output is the converted question text. An appropriate answer is then selected from the FAQ database.

[0862] Step 7:

[0863] The server analyzes the user's emotional state through an emotion recognition engine, adjusts the tone of its response based on that analysis, and sends the response to the terminal. Input is the user's facial expressions and voice data, and output is the adjusted response. The system communicates flexibly according to the user's emotions.

[0864] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0865] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0866] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0867] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0868] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0869] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0870] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0871] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0872] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0873] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0874] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0875] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0876] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0877] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0878] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0879] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0880] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0881] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0882] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0883] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0884] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0885] The following is further disclosed regarding the embodiments described above.

[0886] (Claim 1)

[0887] A means of performing personal authentication based on facial recognition,

[0888] A means of accepting user procedural selections via voice or text,

[0889] A means of automatically generating the documents required for the procedure,

[0890] A means to display the generated document on the user's terminal and enable editing,

[0891] A means of generating answers to user questions using an FAQ database,

[0892] Means of confirming the completion of the procedure,

[0893] A system that includes this.

[0894] (Claim 2)

[0895] The system according to claim 1, which dynamically selects a procedure flow based on the results of facial recognition.

[0896] (Claim 3)

[0897] The system according to claim 1, having voice and text interfaces for guiding the progress of the procedure.

[0898] "Example 1"

[0899] (Claim 1)

[0900] A means of performing personal authentication based on facial recognition,

[0901] A means of accepting user procedural selections via voice or text,

[0902] A means of automatically generating the documents required for the procedure,

[0903] A means for displaying the generated document on an information processing device and enabling editing,

[0904] A means of generating answers to user questions using an information database,

[0905] Means of confirming the completion of the procedure,

[0906] A means of generating appropriate answers to questions using an AI model,

[0907] A system that includes this.

[0908] (Claim 2)

[0909] The system according to claim 1, which dynamically selects a procedure flow based on the results of facial recognition.

[0910] (Claim 3)

[0911] The system according to claim 1, having voice and text interfaces for guiding the progress of the procedure.

[0912] "Application Example 1"

[0913] (Claim 1)

[0914] A means of performing personal authentication based on facial recognition,

[0915] A means of accepting user procedural selections via voice or text,

[0916] A means of automatically generating the documents required for the procedure,

[0917] A means of providing a user interface that allows the generated document to be displayed and edited,

[0918] A means of generating answers to user questions using an information database,

[0919] Means of confirming the completion of the procedure,

[0920] Means for registering or updating payment information,

[0921] A means of obtaining answers to questions using a natural language generation engine,

[0922] A system that includes this.

[0923] (Claim 2)

[0924] The system according to claim 1, which dynamically selects a procedure flow and payment method settings based on the results of facial recognition.

[0925] (Claim 3)

[0926] The system according to claim 1, having voice and text interfaces for guiding the procedure and payment operations.

[0927] "Example 2 of combining an emotion engine"

[0928] (Claim 1)

[0929] A means of identifying individuals based on facial recognition data,

[0930] A means of accepting process selection based on user utterances or input information,

[0931] A means to automatically create the electronic documents required for the process,

[0932] A means of outputting the created electronic document to a display device and making it editable,

[0933] A means of creating responses to user inquiries using a knowledge database,

[0934] A means of confirming that the process has been completed,

[0935] A means of analyzing the user's emotional state and adjusting the guidance content accordingly,

[0936] A system that includes this.

[0937] (Claim 2)

[0938] The system according to claim 1, which dynamically selects process steps based on facial recognition results and provides guidance according to the user's emotional state.

[0939] (Claim 3)

[0940] The system according to claim 1, which has a speech and input information interface for guiding the progress of a process, and adjusts the content and tone of the guidance according to the user's emotions.

[0941] "Application example 2 when combining with an emotional engine"

[0942] (Claim 1)

[0943] Features for individual authentication based on facial recognition,

[0944] A function that accepts procedure selections from the user via voice or text,

[0945] A function to automatically generate the necessary documents for the procedure,

[0946] The generated documents are displayed on the user's terminal and have a function that allows editing,

[0947] A function that generates answers to user questions using an FAQ database,

[0948] A function to confirm the completion of the procedure,

[0949] A function that recognizes the user's emotions and adjusts the procedural guide and responses accordingly,

[0950] Information processing device including

[0951] (Claim 2)

[0952] The information processing device according to claim 1, which dynamically selects a procedure flow based on the results of facial recognition and adjusts the tone and level of detail of the guidance according to the emotional state of the user.

[0953] (Claim 3)

[0954] The information processing device according to claim 1, having voice and text interfaces for guiding the progress of a procedure, and providing responses based on the user's sentiment analysis. [Explanation of Symbols]

[0955] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of performing personal authentication based on facial recognition, A means of accepting user procedural selections via voice or text, A means of automatically generating the documents required for the procedure, A means of providing a user interface that allows the generated document to be displayed and edited, A means of generating answers to user questions using an information database, Means of confirming the completion of the procedure, Means for registering or updating payment information, A means of obtaining answers to questions using a natural language generation engine, A system that includes this.

2. The system according to claim 1, which dynamically selects a procedure flow and payment method settings based on the results of facial recognition.

3. The system according to claim 1, having voice and text interfaces for guiding the procedure and payment operations.