System
A system that processes natural language input to provide personalized assistance and troubleshooting enhances OS usability for all users, including those with visual limitations, by using a generative AI model for intuitive and efficient operation.
Patent Information
- Application Number
- JP2024131434
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2026-02-20
AI Technical Summary
Conventional operating systems (OS) are difficult for users, especially those unfamiliar with technology or with visual limitations, to fully utilize their functions, and there is a lack of effective communication of new features or troubleshooting support.
A system that receives natural language input, analyzes it, and generates responses, including personalized suggestions, application searching, troubleshooting guidance, and feature explanations, using a generative AI model to enhance user interaction and efficiency.
Enables users to operate OS intuitively and efficiently, providing quick troubleshooting and feature adaptation.
Smart Images

Figure 2026028818000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional operating systems (OS) offer numerous functions and icons, but users often fail to fully utilize many of them. Operating an OS is often difficult, especially for users who are unfamiliar with technology or who have visual limitations. This results in a poor user experience, preventing the OS from maximizing its potential utility. Furthermore, when new functions or updates are introduced, there is a lack of a way to quickly communicate this information to users. Furthermore, when troubleshooting, users may not be able to receive appropriate support promptly. There is a need for a system that can resolve these issues and enable users to operate an OS more intuitively and efficiently. [Means for solving the problem]
[0005] This invention solves the above-mentioned problems by providing a system that receives natural language input from a user, analyzes it, generates responses, and displays them. Specifically, the system includes means for acquiring and managing a user profile and providing personalized suggestions to the user. It also includes means for acquiring a user's operation history and configuration information and using it to assist in searching for applications and settings. It also includes means for receiving malfunction reports from users, diagnosing problems, and providing troubleshooting guidance. The system acquires information about new OS features and updates and explains how to use them, enabling users to effectively utilize the new features. It also provides means for improving interaction with the OS through voice and text dialogue for users with visual limitations or language differences. Finally, it includes means for supporting efficient management, switching, and collaboration when operating multiple applications and functions simultaneously. This allows users to operate the OS intuitively and efficiently, enabling quick response to troubleshooting and the introduction of new features.
[0006] "Natural language input" is a format in which a user inputs instructions or questions into a system using everyday language.
[0007] A "user profile" is a database that compiles information about a user, including past operation history and individual setting information.
[0008] A "generative AI model" is an artificial intelligence that uses machine learning algorithms to analyze user input and generate appropriate responses.
[0009] A "fault report" is an act by a user reporting a technical problem or trouble to the system.
[0010] "Troubleshooting" is the process of diagnosing a problem and providing the user with a solution.
[0011] "Personalized suggestions" refers to providing individually customized suggestions and guidance based on a user's profile and past operation history.
[0012] "Searching for applications and settings" refers to the navigation and search functionality provided by the system to help users find specific applications and settings.
[0013] "Response generation" means that the system automatically generates appropriate dialogue and guidance in response to the user's natural language input.
[0014] "Introduction and explanation of new features" means that the system acquires information about new features and updates, and notifies and explains their contents and how to use them to users.
[0015] "Visual constraints" refers to a situation in which a user has a visual limitation or disability that makes it difficult to use a typical visual interface.
[0016] "Operation history" refers to a record of a series of operations or actions that a user has performed in the past.
[0017] "Update information" refers to information about changes and new features when the software is updated to a new version. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] As an embodiment of the present invention, we provide a chatbot system that supports the use of operating systems (OS) for PCs and smartphones. The specific operation method and processing content of this system are described below.
[0040] First, the system includes means for receiving natural language input from a user, means for analyzing the received input, means for generating a response based on the analysis result, and means for displaying the response to the user.
[0041] Program processing overview
[0042] 1. Initialization and Setup
[0043] The server hosts the chatbot application, loads the generative AI model, and exposes API endpoints for interacting with the OS and users, starting the chatbot.
[0044] The device launches a dedicated chatbot app built into the OS, which displays a user interface (UI) and allows the user to enter a message.
[0045] The user accesses the chatbot application and starts a dialogue. The chatbot receives the natural language input from the user in text format.
[0046] 2. User Authentication and Profiling
[0047] A user logs in to an application. The server checks the user information and retrieves the user profile. The retrieved user profile includes past operation history and setting information. This allows personalized suggestions to be made to the user.
[0048] The device will prepare individually customized suggestions and guides based on the user's settings and past operation history.
[0049] 3. Natural Language Understanding and Dialogue Generation
[0050] Users enter questions or instructions into a chat box (e.g., "I want to install a new app"). The device sends this input to the server, which uses generative AI models to parse the natural language and generate an appropriate response.
[0051] The generated response is sent to the terminal and displayed on the user interface, allowing the user to intuitively understand the procedures for operating the OS.
[0052] 4. Function execution and operation support
[0053] Based on the response from the server, the device will take specific action. For example, if the user inputs "I want to install a new app," the server will search for the appropriate application and generate a response guiding the user through the installation process. The device will then display this response and guide them through the necessary operations.
[0054] The user follows the instructions to perform the action (e.g., clicks the install button). The device tracks the operation status in real time and displays additional guidance as needed.
[0055] 5. Troubleshooting
[0056] When a user reports a specific problem (e.g., "I can't connect to Wi-Fi"), the server analyzes the problem and identifies possible causes. The generated troubleshooting guidance is sent to the device as specific solution steps. The device displays the solution steps on the user interface and provides detailed instructions to the user.
[0057] Specific examples
[0058] New user onboarding
[0059] User: Starts up a newly purchased PC and uses the chatbot for the first time.
[0060] Device: Display "Nice to meet you. Let's go over the basic settings together. First, let's set the language."
[0061] User: Type "Please set to Japanese."
[0062] Terminal: Send this input to the server.
[0063] Server: Analyzes the message and generates instructions to change the language setting to Japanese.
[0064] Device: Change the language setting to Japanese and notify the user that "Settings are complete."
[0065] Introducing and explaining new features
[0066] User: After an update, launch the chatbot and type, "Tell me about the new features."
[0067] Server: Searches for information about new features and generates a summary of them.
[0068] On your device: "The new update includes a refreshed design, battery optimizations, and new security features."
[0069] Fault diagnosis and troubleshooting
[0070] User: Reports "No sound."
[0071] Server: Analyzes the problem, lists possible causes, and generates solutions.
[0072] Device: "Possible causes of no sound include the volume setting, muting, or speaker failure. First, check the volume setting." is displayed.
[0073] Users: Follow the instructions and double-check your volume settings.
[0074] In this way, a dedicated chatbot can assist users in various situations, making operating the OS more intuitive and efficient. This system allows users to effectively use the OS's functions and settings, and quickly respond to any problems that arise.
[0075] The processing flow will be explained below.
[0076] Step 1:
[0077] The server launches the chatbot application and loads the generative AI model, which is a machine learning model for natural language analysis and response generation. The server also opens an API endpoint and prepares to accept connection requests from devices.
[0078] Step 2:
[0079] The device installs and launches the chatbot application. The user interface (UI) is displayed, allowing the user to enter a message. The initial screen displays, "Nice to meet you. Welcome to the chatbot. How can I help you?"
[0080] Step 3:
[0081] The user enters login information (e.g., user ID and password) into the UI and clicks the "Login" button, which sends a login request from the terminal to the server.
[0082] Step 4:
[0083] The server compares the received login information with a database and performs authentication. If authentication is successful, the server obtains the user profile and sends it to the terminal. The user profile includes past operation history and individual setting information.
[0084] Step 5:
[0085] The device will then display a personalized UI based on the received user profile, such as shortcuts to frequently used functions and recently accessed applications.
[0086] Step 6:
[0087] The user types a question or instruction into the chat box (e.g., "I want to install a new app"), and the message is sent directly to the server.
[0088] Step 7:
[0089] The server receives the user's input message and uses a generative AI model to perform natural language analysis, identifying the user's intent and request (e.g., "I want to know the steps to install a new app").
[0090] Step 8:
[0091] The server generates an appropriate response based on the analysis results, such as a command to open the application store or guidance on installation procedures.
[0092] Step 9:
[0093] The generated response is sent from the server to the device, and may include specific instructions or options.
[0094] Step 10:
[0095] The device displays the received response on the user interface, for example, a confirmation message such as "Do you want to open the application store?"
[0096] Step 11:
[0097] The user follows the displayed instructions (e.g., answers "yes" and opens the application store), and the user's selection is sent back to the server.
[0098] Step 12:
[0099] Based on the user's selection, the server determines the next action and sends the necessary instructions to the terminal, such as to display search results for a particular application.
[0100] Step 13:
[0101] The device updates its user interface based on the instructions received from the server, for example by displaying a list of search results and allowing the user to select the app they want to install.
[0102] Step 14:
[0103] The user selects the app they want to install and sends an installation request to the server again.
[0104] Step 15:
[0105] The server confirms the installation request, initiates the necessary installation process, and notifies the device of the installation progress.
[0106] Step 16:
[0107] The terminal displays the installation progress to the user and notifies the user when the installation is complete.
[0108] Step 17:
[0109] The user launches the application after the installation is complete and begins using it.
[0110] Through this process, users can interact with the OS in natural language and smoothly perform the required actions.
[0111] Example 1
[0112] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0113] In recent years, the use of information devices such as PCs and smartphones has continued to expand, resulting in a demand for support regarding device operation and settings. However, conventional support systems have difficulty providing personalized support based on individual users' needs and operation history, limiting the improvement of usability. Furthermore, when it comes to troubleshooting, it is difficult to provide a prompt and accurate response to specific problems users face.
[0114] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0115] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, means for generating a corresponding response based on the analysis, means for displaying the response to the user, means for opening an API endpoint for communication with an OS and a user, means for loading a generative AI model, means for analyzing the user's input and generating an appropriate response, means for sending the generated response to a terminal, means for customizing a user interface, means for tracking user operations in real time and displaying additional guidance, and means for analyzing problems, identifying possible causes, and generating solutions. This enables personalized support based on the user's individual needs and operation history, and also enables quick and accurate troubleshooting.
[0116] "Natural language input from a user" refers to a message expressed in everyday language that is input by a user in the form of text, voice, or the like.
[0117] "Means for analyzing natural language input" refers to an algorithm or system that understands the natural language message received from the user and analyzes its grammatical structure and meaning.
[0118] "Means for generating a corresponding response" refers to a system or algorithm for generating an appropriate reply or instruction based on the results of natural language analysis.
[0119] "Means for displaying a response to a user" refers to an interface or device for displaying the generated response so that the user can visually confirm it.
[0120] "Means for opening API endpoints" refers to a system that provides the protocols and interfaces necessary to communicate with information terminals and operating systems.
[0121] "Means for loading a generative AI model" refers to the process or algorithm used to load a generative model into memory or a processor and make it executable.
[0122] "Means for analyzing user input and generating an appropriate response" refers to a process or system that analyzes user input and generates an optimal response based on the results.
[0123] The "means for transmitting the generated response to the terminal" refers to a system for sending the generated response data to the terminal via a network.
[0124] "Means for customizing a user interface" refers to a method or system for changing the design and functionality of an interface based on user settings and operation history.
[0125] "Means for tracking user operations in real time and displaying additional guidance" refers to a system that monitors the operations performed by a user in real time and provides next steps or additional instructions as needed.
[0126] "Means for analyzing problems, identifying possible causes, and generating solutions" refers to the systems and processes used to analyze user-reported problems, identify their causes, and generate appropriate solutions.
[0127] This invention relates to a chatbot system that supports the use of operating systems (OS) for PCs and smartphones. Specifically, it is a system that uses a generative AI model to analyze natural language input from users and generate appropriate responses.
[0128] This system mainly consists of three elements: a server, a terminal, and a user.
[0129] server
[0130] The server hosts the chatbot application and loads the generative AI model (e.g., GPT-4). The server exposes API endpoints for communication with the OS and users, and receives and analyzes user input. The server has the following capabilities:
[0131] A means of receiving natural language input from the user
[0132] A means of parsing natural language input
[0133] A means of generating a corresponding response based on the analysis results
[0134] A means of sending the generated response to the terminal
[0135] A means of analyzing problems, identifying possible causes, and generating solutions
[0136] Specifically, the server performs the following process.
[0137] 1. When receiving input from the user, retrieve the data through the appropriate API endpoint.
[0138] 2. The received input data is analyzed using a generative AI model, which performs tokenization and grammar analysis to understand meaning.
[0139] 3. Based on the analysis results, the server generates an appropriate response, referencing the user profile and past operation history to provide personalized suggestions.
[0140] 4. The generated response is sent to the device through the API endpoint.
[0141] Terminal
[0142] The device runs a dedicated chatbot application built into the operating system. The device is equipped with the following means:
[0143] A way to customize the user interface
[0144] A means to track user actions in real time and display additional guidance
[0145] Specifically, the terminal operates as follows.
[0146] 1. Launch the chatbot application and display a user interface that allows the user to enter a message.
[0147] 2. If there is any input from the user, send it to the server.
[0148] 3. Display the response received from the server in the user interface, providing interactive instructions as needed.
[0149] 4. Track user actions in real time and provide appropriate guidance when information is missing or additional guidance is needed.
[0150] User
[0151] Users access the chatbot application using a computer or smartphone and receive support through natural language dialogue. Users mainly perform the following operations:
[0152] 1. Log in to the application and provide the required authentication information.
[0153] 2. Enter questions or instructions in the chat box and receive responses from the server.
[0154] 3. Follow the displayed instructions and guides to operate the OS.
[0155] Specific examples
[0156] For example, when a user who has just purchased a new PC performs initial setup, the following dialogue may take place:
[0157] User: "I just started up my new PC. What do I do now?"
[0158] Server: "Nice to meet you. Let's go over some basic setup together. First, let's set up the language."
[0159] User: "Please set it to Japanese."
[0160] Server: "The language setting has been changed to Japanese. Next, set the time zone."
[0161] Also, if the user wants to install a new application, the prompt sentence is as follows:
[0162] Prompt for "I want to install a new app":
[0163] Input to the server: "The user wants to install a new app. Please guide them through the installation process."
[0164] Example response: "Find the right application and guide me through the installation process."
[0165] This invention allows users to intuitively and efficiently use OS functions and settings, and also allows them to receive prompt and accurate support when problems occur.
[0166] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0167] Step 1: Initialization and Setup
[0168] server:
[0169] The server hosts the chatbot application, loads the generative AI model (e.g., GPT-4), opens an API endpoint, and prepares to communicate with the OS and the user. Specific operations include the following steps:
[0170] The server reads specific configuration files and database connection information.
[0171] The server loads the generative AI model into memory and makes it executable.
[0172] Device:
[0173] The device launches a dedicated chatbot app and displays a user interface (UI), allowing the user to enter a message. Specifically, the process involves the following steps:
[0174] The device prepares the UI for display, rendering text boxes, submit buttons, etc.
[0175] The terminal checks the connection to the server and whether communication is possible.
[0176] User:
[0177] A user accesses a chatbot application and starts a conversation. The specific operations include the following steps:
[0178] The user checks the application's UI and enters an initial message such as "Hello."
[0179] Step 2: User authentication and profiling
[0180] User:
[0181] A user logs into an application, typically by entering a username and password, and sometimes by using biometric authentication such as facial recognition or fingerprint authentication.
[0182] server:
[0183] The server checks the user's authentication information against the database and retrieves the user profile. Specifically, the following process is performed:
[0184] The server takes the entered username and password and issues an SQL query to the database to verify it.
[0185] If the authentication is successful, the server obtains the user profile (including operation history and setting information) and stores it in memory.
[0186] Device:
[0187] The device customizes the UI based on the user profile information received from the server. Specifically, the following process is performed:
[0188] The device changes the UI style based on the user's language and theme settings.
[0189] Step 3: Natural Language Understanding and Dialogue Generation
[0190] User:
[0191] The user types a question or instruction into the chat box, for example, "I want to install a new app."
[0192] Device:
[0193] The terminal sends the user's input to the server. The specific procedure is as follows:
[0194] The device receives the input as text data and sends an HTTP POST request through the API endpoint.
[0195] server:
[0196] The server analyzes the received user input using a generative AI model. Specifically, the following data processing and calculations are performed:
[0197] The server tokenizes the input data, performs grammatical analysis, and analyzes it to understand its meaning.
[0198] Based on the analysis results, an appropriate response is generated.
[0199] Device:
[0200] The generated response is sent to the terminal and displayed on the user interface. Specifically, the process includes the following steps:
[0201] The terminal acquires the response data received from the server and displays it in the text box.
[0202] Step 4: Implementing functions and assisting operations
[0203] server:
[0204] The server then generates a response that recommends a specific action, such as instructions for installing a new app.
[0205] Device:
[0206] The terminal guides the user according to the procedure from the server. Specifically, the terminal performs the following procedure.
[0207] The device will display instructions such as "Click the install button."
[0208] User:
[0209] The user follows the instructions displayed on the device, for example, "click the install button."
[0210] Step 5: Troubleshooting
[0211] User:
[0212] The user reports a specific problem, for example, "I can't connect to Wi-Fi."
[0213] server:
[0214] The server analyzes the problem, identifies possible causes, and generates solutions. Specifically, the following data processing and calculations are performed:
[0215] The server references the data for troubleshooting and identifies the cause through natural language analysis.
[0216] The server generates solutions and organizes them into guidance.
[0217] Device:
[0218] The generated troubleshooting guidance is displayed in the user interface, specifically by following the steps below.
[0219] The device will display instructions such as "First, please check your volume settings."
[0220] This system allows users to operate the OS and solve problems intuitively and efficiently.
[0221] (Application example 1)
[0222] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0223] Conventional systems make it difficult for users to operate and configure the system in natural language, and they lack sufficient support for security-related settings and troubleshooting. Furthermore, they do not provide personalized suggestions that effectively utilize user profiles and operation histories, making it difficult for users to operate the system intuitively and efficiently.
[0224] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0225] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, and means for generating a corresponding response based on the analysis, thereby enabling the user to perform security settings and other operations in natural language.
[0226] Furthermore, the server includes means for performing specific actions based on the generated responses and means for analyzing problems and generating solutions during troubleshooting, thereby enabling the user to obtain specific troubleshooting and operation guidance.
[0227] Furthermore, by including a means for acquiring and managing a user profile, a means for providing personalized suggestions to the user based on the profile, and a means for adjusting security settings based on the user's operation history and setting information, intuitive and individually optimized operation becomes possible, allowing the user to efficiently manage security settings.
[0228] "Natural language input" is an input format in which a user gives instructions or asks questions to a system in everyday language.
[0229] "Analysis" refers to the process of converting input natural language into a form that a machine can understand and interpreting its meaning.
[0230] A "response" is a reply or instruction from the system that is generated based on the parsed natural language input.
[0231] A "user profile" is a data set that collects information about a user, including individual settings and past operation history.
[0232] "Personalized suggestions" are suggestions that individually provide users with information and guidance that is most suitable for them based on their user profile.
[0233] The "operation history" is a record of a series of operations and setting changes that the user has performed in the past.
[0234] A "specific action" is a specific operation or process that the system actually performs based on a user's instruction.
[0235] "Troubleshooting" is the process of identifying the cause of a problem and providing a solution.
[0236] A "generative AI model" is an artificial intelligence technology that uses machine learning to understand natural language and generate appropriate responses.
[0237] "Security settings" refers to a set of settings and policies that ensure the security of a device or system.
[0238] As an embodiment of the present invention, we will explain a chatbot system in which a server and a terminal cooperate to assist users in security configuration and troubleshooting. This system utilizes a generative AI model to provide appropriate responses to security-related instructions and questions entered by the user in natural language.
[0239] Programs and hardware / software used
[0240] An API (specifically, the OpenAI API) that runs the generative AI model is installed on the server, which analyzes input from users and generates responses. A smartphone is used as the terminal, and a chatbot app is installed on this smartphone. Users interact with the system through this app.
[0241] Natural language input processing begins when a user enters instructions or questions in natural language into the chatbot app. This input is sent to the server as text data, where it is analyzed by the generative AI model. Based on the results of the analysis, the server generates an appropriate response and sends it back to the device.
[0242] User profile management is achieved by applications sending user information to a server, which then stores the user's operation history and settings information and generates personalized suggestions and actions based on this information, which can be useful when adjusting security settings.
[0243] For example, if a user types "check Wi-Fi settings," the device will execute a script based on the generated response to help check the actual Wi-Fi settings, allowing users to easily troubleshoot security settings and resolve issues.
[0244] When troubleshooting, users report their issues in natural language and their input is sent to the server, which uses generative AI models to analyze the cause of the problem and generate a solution as a response. For example, if a user reports "no sound," specific steps such as "check your volume settings" are provided.
[0245] Specific examples
[0246] User: "Check the security status of your device"
[0247] Chatbot: "We're running a security check on your device. Please wait a moment... All security settings are OK."
[0248] Example prompt sentence:
[0249] User input: Check the security status of your device
[0250] Security Chatbot Response: We're going to perform a security check on your device. Please wait... All your security settings are fine.
[0251] In this way, users can easily perform security-related operations in natural language through the chatbot and receive appropriate troubleshooting guidance when problems arise, enabling users to manage their security settings intuitively and efficiently.
[0252] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0253] Step 1:
[0254] The device launches the chatbot app, and the user inputs instructions or questions in natural language. The input text data is sent to the server via an API call. The server receives the text data and prepares it for analysis by the generative AI model.
[0255] Input: User's natural language input (text data)
[0256] Output: Text data sent to the server
[0257] Step 2:
[0258] The server passes the received natural language input to the generative AI model, which begins the analysis process. The generative AI model analyzes the input natural language, understands its meaning, and generates an appropriate response. The analysis uses the OpenAI API.
[0259] Input: Text data sent to the server
[0260] Output: Analysis results from the generative AI model
[0261] Step 3:
[0262] Based on the generated analysis, the server generates a specific response to the user's instruction, such as "check your Wi-Fi settings" or any other troubleshooting steps required.
[0263] Input: Analysis results from generative AI model
[0264] Output: Response data to the user
[0265] Step 4:
[0266] The server sends the generated response data to the terminal, which receives the response and displays it on the user interface within the chatbot app. The user can then view it and take the next action.
[0267] Input: Response data to the user
[0268] Output: Response data sent to the terminal
[0269] Step 5:
[0270] The user takes action based on the generated response. For example, if the instruction is "Check your Wi-Fi settings," the user opens the Wi-Fi settings screen and adjusts the settings. After completing this action, the user again provides feedback to the chatbot.
[0271] Input: Response data displayed to the user
[0272] Output: User's actual actions
[0273] Step 6:
[0274] If the user wants to input additional instructions to resolve the issue, the process repeats from step 1. For example, if the user inputs "Please also check the volume settings," the server will analyze it again and return the appropriate instructions.
[0275] Input: Additional user instructions (natural language input)
[0276] Output: The analysis process is restarted.
[0277] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0278] As an embodiment of the present invention, we provide a chatbot system that supports the use of operating systems (OS) for PCs and smartphones. In particular, we will explain an embodiment that incorporates an engine that recognizes the user's emotions.
[0279] First, the system has the means to receive natural language input from users, parse it, generate and display a response, and, combined with an emotion engine, recognizes the user's emotional state and adapts responses based on that, achieving a high level of personalization.
[0280] Program processing overview
[0281] 1. Initialization and Setup
[0282] The server hosts the chatbot application and loads the generative AI model, a machine learning model for natural language analysis and response generation, and the emotion engine, a model for recognizing user emotions.
[0283] The chatbot application is installed on the device and launched. The user interface (UI) is displayed, allowing the user to enter a message. The initial screen displays, "Hello. Is there anything I can help you with?"
[0284] The user accesses the chatbot application and starts a dialogue, receiving the user's natural language or voice input in text format.
[0285] 2. User Authentication and Profiling
[0286] A user logs in to an application. The server checks the user information and retrieves the user profile. The retrieved user profile includes past operation history and setting information.
[0287] The device displays a personalized UI based on the user's settings and past operation history, allowing it to provide more appropriate suggestions and guidance to the user.
[0288] 3. Natural Language Understanding and Emotion Recognition
[0289] Users enter questions or instructions into the chat box (e.g., "I want to install a new app"). The system also analyzes the user's facial expressions and tone of voice using an emotion engine. The device then sends these inputs to the server.
[0290] The server uses a generative AI model to analyze natural language and an emotion engine to recognize the user's emotions. The analysis results identify the user's intent, request, and emotional state (e.g., "I want to know the steps to install a new app. The user is feeling a bit anxious and worried.").
[0291] 4. Response generation and emotional adaptation
[0292] Based on the analysis and emotional state, the server generates an appropriate response, such as a command to open the application store or friendly, gentle installation instructions.
[0293] The generated response is sent from the server to the device, and may include specific instructions or an emotionally sensitive tone.
[0294] The device displays the received response on the user interface, for example, a message such as "Open the application store and install the desired app. If you have any questions, please feel free to ask."
[0295] 5. Troubleshooting
[0296] When a user reports a specific problem (e.g., "I can't connect to Wi-Fi"), the server analyzes the problem and evaluates the user's emotional state using an emotion engine. The generated troubleshooting guidance is sent to the device as specific solution steps. The device then displays the solution steps on the user interface and provides emotion-sensitive instructions.
[0297] Specific examples
[0298] New user onboarding
[0299] User: Starts up a newly purchased PC and uses the chatbot for the first time.
[0300] Device: Display "Hello. Let's go through the basic setup together. If you have any questions, please feel free to ask."
[0301] User: Type "Please set to Japanese."
[0302] Terminal: Send this input to the server.
[0303] Server: Parses the message and generates instructions to change the language setting to Japanese. Recognizes that the user is feeling a bit anxious and generates a response in a friendly tone.
[0304] Device: Change the language setting to Japanese and notify the user, "Setup complete. You can now use Japanese. Is there anything else we can help you with?"
[0305] Introducing and explaining new features
[0306] User: After an update, launch the chatbot and type, "Tell me about the new features."
[0307] Server: Searches for information about new features and generates summaries of them, recognizing the user's emotional state and crafting interesting descriptions.
[0308] On your device: "Our new update features a fresh design, battery optimizations, and new security features."
[0309] Fault diagnosis and troubleshooting
[0310] User: Reports "No sound."
[0311] Server: Analyzes the problem and recognizes the user's emotional state using an emotion engine. Lists possible causes and generates solutions.
[0312] Device: "Possible reasons for the lack of sound include the volume setting, muting, or a faulty speaker. First, check the volume setting. If you have any questions, please let us know."
[0313] User: Follow the instructions and check the volume settings again. In this way, a system incorporating an emotion engine allows users to receive responses that take their emotions into consideration, allowing them to operate the OS with greater peace of mind.
[0314] The processing flow will be explained below.
[0315] Step 1:
[0316] The server launches the chatbot application and loads the generative AI model and emotion engine. The generative AI model is a machine learning model for natural language analysis and response generation, and the emotion engine is a model for recognizing user emotions. The server also opens an API endpoint and prepares to accept connection requests from devices.
[0317] Step 2:
[0318] The chatbot application is installed on the device and launched. The user interface (UI) is displayed, allowing the user to enter a message. The initial screen displays, "Hello. Is there anything I can help you with?"
[0319] Step 3:
[0320] The user enters login information (e.g., user ID and password) into the UI and clicks the "Login" button, which sends a login request from the terminal to the server.
[0321] Step 4:
[0322] The server compares the received login information with a database and performs authentication. If authentication is successful, the server obtains the user profile and sends it to the terminal. The user profile includes past operation history and individual setting information.
[0323] Step 5:
[0324] The device will then display a personalized UI based on the received user profile, such as shortcuts to frequently used functions and recently accessed applications.
[0325] Step 6:
[0326] Users enter questions or instructions into the chat box (e.g., "I want to install a new app"). The entered message is sent directly to the server. The system also analyzes the user's facial expressions and tone of voice using an emotion engine.
[0327] Step 7:
[0328] The server receives the user's input message and uses a generative AI model to perform natural language analysis, identifying the user's intent and request (e.g., "I want to know the steps to install a new app").
[0329] Step 8:
[0330] The server uses an emotion engine to recognize the user's emotional state, for example, to determine whether the user is feeling anxious or worried.
[0331] Step 9:
[0332] The server generates an appropriate response based on the analysis and emotional state, for example, "Do you want to open the app store?" in a calm and reassuring tone.
[0333] Step 10:
[0334] The generated response is sent from the server to the device, and may include specific instructions or an emotionally sensitive tone.
[0335] Step 11:
[0336] The device displays the received response on the user interface, for example, "Open the application store and install the desired app. If you have any questions, please feel free to ask."
[0337] Step 12:
[0338] The user follows the displayed instructions (e.g., answers "yes" and opens the application store), and the user's selection is sent back to the server.
[0339] Step 13:
[0340] Based on the user's selection, the server determines the next action and sends the necessary instructions to the terminal, such as to display search results for a particular application.
[0341] Step 14:
[0342] The device updates its user interface based on the instructions received from the server, for example by displaying a list of search results and allowing the user to select the app they want to install.
[0343] Step 15:
[0344] The user selects the app they want to install and sends an installation request to the server again.
[0345] Step 16:
[0346] The server confirms the installation request, initiates the necessary installation process, and notifies the device of the installation progress.
[0347] Step 17:
[0348] The terminal displays the installation progress to the user and notifies the user when the installation is complete.
[0349] Step 18:
[0350] Once the installation is complete, the user launches the application and begins using it. Through this process, the user can interact with the OS based on natural language and emotions to smoothly perform the required actions.
[0351] Example 2
[0352] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0353] Conventional chatbot systems have difficulty generating appropriate responses to users' natural language input, and in particular, they are unable to recognize emotions and reflect them in responses, which means they are unable to provide a satisfying conversational experience for users.In addition, they lack personalized suggestions that utilize the user's profile information and past operation history, making it difficult to achieve user-friendly operation.
[0354] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0355] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, means for generating a corresponding response based on the analysis, means for displaying the response to the user, means for analyzing natural language using a generative AI model, means for identifying the user's emotional state using an emotion recognition engine, means for generating an appropriate response based on the analysis result and the emotional state, and means for displaying the response in a tone that takes the user's emotions into consideration. This makes it possible to provide more personalized suggestions and guidance while taking the user's emotions into consideration, thereby achieving a highly satisfying dialogue experience.
[0356] "Natural language input" refers to text or voice input entered by a user in the form of everyday spoken language.
[0357] A "generative AI model" refers to an algorithmic model that uses machine learning technology to analyze natural language and generate responses.
[0358] "Emotion recognition engine" refers to a software module for analyzing and identifying an emotional state from a user's input.
[0359] A "user profile" refers to a collection of data including a user's personal information, past operation history, and setting information.
[0360] "Personalization" refers to providing customized suggestions and responses based on individual user characteristics and past behavior.
[0361] "Operation history" refers to a record of operations that a user has performed in the past.
[0362] "Setting information" refers to information about various settings that a user makes when using an application or system.
[0363] A "server" refers to a computer or software system used to receive requests from users and process or respond to data.
[0364] "Terminal" refers to the device that the user actually operates (e.g., a PC or smartphone).
[0365] "User interface" refers to the screen display and input means that allow a user to interact with a system.
[0366] "Login" refers to the authentication procedure required for a user to access a system.
[0367] "Analysis" refers to the processing of data to understand a user's natural language input and emotional state.
[0368] A "response" refers to a reply or instruction that a system generates and provides in response to a user's input.
[0369] As an embodiment of the present invention, we provide a chatbot system that supports the use of operating systems (OS) for personal computers and smartphones. This system incorporates a generative AI model and an emotion recognition engine.
[0370] First, the server hosts the chatbot application and loads a generative AI model (e.g., GPT-3.5) and an emotion recognition engine (e.g., a popular emotion recognition API). The generative AI model is used to parse natural language and generate responses, and the emotion recognition engine is used to recognize the user's emotional state. These models and engines are initialized by the server by loading the API key and model data.
[0371] The device then allows the user to install and launch the chatbot application. When the user launches the application, a user interface appears, displaying a welcome message that reads, "Hello, how can I help you?"
[0372] The user accesses the chatbot application and starts a dialogue using natural language or voice. The input content is received in text format and sent from the terminal to the server. If the user inputs voice, the voice is converted into text and sent to the server.
[0373] When a user logs in to an application, the server checks the user information and obtains profile data (past operation history, setting information, etc.). Based on the obtained profile data, a personalized UI is displayed to the user. For example, frequently used functions can be displayed as shortcuts based on the past operation history.
[0374] The server analyzes the natural language input from the user using a generative AI model to identify the user's intentions and requests. It also uses an emotion recognition engine to recognize the user's emotional state. For example, if a user inputs "I want to install a new app," the analysis results might determine that "I want to know the steps to install a new app, and I'm feeling a bit impatient and anxious."
[0375] Based on the analysis results and the emotional state, the server uses a generative AI model to generate an appropriate response, such as "Please visit the application store and install the app you want. If you have any questions, please feel free to ask." This response is generated in an emotionally sensitive tone.
[0376] The generated response is sent to the device and displayed in the user interface. The displayed response is provided in an emotionally sensitive tone, allowing the user to proceed with the operation with confidence.
[0377] When a user reports a specific problem, for example, "I can't connect to Wi-Fi," the server analyzes the problem and uses an emotion recognition engine to assess the user's emotional state. Based on the analysis, specific solutions are generated and sent to the device. For example, a message might read, "Possible causes of the lack of audio include volume settings, muting, or a faulty speaker. First, check your volume settings."
[0378] Specific examples
[0379] New user onboarding
[0380] User: Starts up a newly purchased PC and uses the chatbot for the first time.
[0381] Device: Display "Hello. Let's go through the basic setup together. If you have any questions, please feel free to ask."
[0382] User: Type "Please set to Japanese."
[0383] Terminal: Send this input to the server.
[0384] Server: Parses the message and generates instructions to change the language setting to Japanese. Recognizes that the user is feeling a bit anxious and generates a response in a friendly tone.
[0385] Device: Change the language setting to Japanese and notify the user, "Setup complete. You can now use Japanese. Is there anything else we can help you with?"
[0386] Introducing and explaining new features
[0387] User: After an update, launch the chatbot and type, "Tell me about the new features."
[0388] Server: Searches for information about new features and generates summaries of them, recognizing the user's emotional state and crafting interesting descriptions.
[0389] On your device: "Our new update features a fresh design, battery optimizations, and new security features."
[0390] Fault diagnosis and troubleshooting
[0391] User: Reports "No sound."
[0392] Server: Analyzes the problem and recognizes the user's emotional state using an emotion recognition engine. Lists possible causes and generates solutions.
[0393] Device: "Possible reasons for the lack of sound include the volume setting, muting, or a faulty speaker. First, check the volume setting. If you have any questions, please let us know."
[0394] In this way, a system incorporating an emotion recognition engine allows users to receive responses that take their emotions into consideration, allowing them to operate the OS with greater peace of mind.
[0395] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0396] Step 1:
[0397] The server hosts the chatbot application and loads the generative AI model (e.g., GPT-3.5) and emotion recognition engine (e.g., a popular emotion recognition API). Specifically, it loads the API key and model data to initialize the model. The input at this stage is the required API key and model data, and the output is a ready-to-use generative AI model and emotion recognition engine.
[0398] Step 2:
[0399] When a user installs and launches a chatbot application, the device loads the user interface. The initial screen displays "Hello, is there anything I can help you with?" When a user launches the application, the input is the application launch command, and the output is a user interface displaying the message.
[0400] Step 3:
[0401] The user inputs questions or instructions to the chatbot using natural language or voice. For example, they might input "I want to install a new app." This input is received by the device in text format. If the input is voice, the voice is converted to text. At this stage, the input is the user's natural language or voice input, and the output is text data.
[0402] Step 4:
[0403] The device sends the received text data to the server. It also sends text data converted from speech. Specifically, it sends the text data to the server via an HTTP request. At this stage, the input is user data in text format, and the output is the data sent to the server.
[0404] Step 5:
[0405] The server analyzes the received text data using a generative AI model to identify the user's intentions and requests. It also uses an emotion recognition engine to recognize the user's emotional state. For example, if a user inputs "I want to install a new app," the analysis results might determine that "I want to know the steps to install a new app. The user is feeling a bit impatient and anxious." The input at this stage is text data, and the output is the analysis results regarding the user's intentions and emotional state.
[0406] Step 6:
[0407] The server uses a generative AI model to generate an appropriate response based on the analysis results and the emotional state. For example, it generates a message such as, "Open the application store and install the app you want. If you have any questions, please feel free to ask." This response is generated in an emotionally sensitive tone. The input at this stage is the analysis results and the emotional state data, and the output is an appropriate response message.
[0408] Step 7:
[0409] The server sends the generated response message to the device. The device receives the response and displays it on its user interface. The displayed message contains specific instructions such as "Open the application store and install the desired app." The input at this stage is the response message, and the output is the message displayed on the user interface.
[0410] Step 8:
[0411] When a user reports a specific problem, for example, "I can't connect to Wi-Fi," the server analyzes the problem and uses an emotion recognition engine to assess the user's emotional state. The analysis results in a list of possible causes and generates a solution. For example, a message might read, "Possible causes of the no sound include the volume setting, muting, or a faulty speaker. Please check the volume setting first." The input at this stage is text data about the specific problem, and the output is the problem analysis and a solution message.
[0412] Step 9:
[0413] The terminal displays the solution message received from the server on the user interface. For example, it displays "Please check your volume settings." This message is expressed in an emotionally sensitive tone. The input at this stage is the solution message, and the output is the solution message displayed on the user interface.
[0414] This allows users to receive emotionally sensitive responses and operate with peace of mind.The system improves the user experience by combining natural language input analysis and emotion recognition.
[0415] (Application example 2)
[0416] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0417] In factory work environments, workers are required to be able to operate robots and troubleshoot them quickly and reliably. However, conventional systems have difficulty responding to the user's emotional state and real-time work environment, which often causes stress and confusion for workers. This can hinder efficient work execution and reduce safety.
[0418] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0419] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, means for generating a corresponding response, means for recognizing the user's emotional state, and means for adapting the response based on the user's emotional state. This enables generation of a response that takes the user's emotions into consideration. The server also includes means for acquiring the user's operation history and setting information, means for providing information related to the work environment in real time, and means for providing appropriate operation guidance to the user. This enables appropriate guidance based on real-time information about the work environment, allowing work to be performed safely and efficiently.
[0420] "Natural language input" refers to a language that is input by a user in the form of text, speech, or other language that is commonly used by humans.
[0421] "Analysis" refers to the syntactic and semantic processing and interpretation of input natural language.
[0422] A "response" is a response to a question or instruction that is generated based on the results of the analysis.
[0423] "Display" refers to providing the generated response to the user visually or audibly.
[0424] "Emotional state" refers to the user's psychological and emotional state as determined by their tone of voice and facial expression.
[0425] A "profile" is a collection of data about a user, including the user's personal information, operation history, setting information, and the like.
[0426] "Personalization" refers to providing individualized offers and information based on a user's profile information.
[0427] An "operation history" is a record of operations and actions that a user has performed in the past.
[0428] "Setting information" refers to setting items and parameters that are individually set by the user for the system or application.
[0429] "Work environment information" is information relating to the user's real-time work location and situation.
[0430] "Operational guidance" refers to procedures or instructions provided to a user to perform a specific task or operation.
[0431] This embodiment of the present invention is a chatbot system with advanced response capabilities incorporating user emotion recognition. It is particularly intended for supporting robotic operation in factory work environments. The system analyzes the user's natural language input, generates corresponding responses, and displays those responses. It also provides appropriate operational guidance in real time based on the user profile and work environment information.
[0432] 1. Hardware and Software Used
[0433] This system uses the following hardware and software:
[0434] Hardware: Smartphones, smart glasses, head-mounted displays
[0435] software:
[0436] Python: To execute the main logic of the program
[0437] OpenAI API: Natural Language Analysis and Response Generation
[0438] Emotion Recognition API: To recognize user emotions
[0439] 2. Program processing content
[0440] The server uses a generative AI model to analyze natural language and generate appropriate responses to user input. Furthermore, it uses the Emotion Recognition API to determine the user's emotional state and adapts responses based on that emotion. Specific responses include:
[0441] 3. System Operation Example
[0442] Operation Guidance
[0443] When a user inquires about how to operate the robot, the server analyzes the information and provides appropriate operating instructions. For example, if a user asks, "How do I install a new part?", the server generates a response like this: "The installation steps are as follows: First, turn off the power and check the safety devices. Then install the part. Please let us know if you have any questions."
[0444] troubleshooting
[0445] When a user reports a problem with the robot, the server analyzes the cause and provides instructions based on emotion recognition to reduce stress and confusion. For example, if a user types, "My robot has stopped working, what should I do?", the server generates a response such as, "Please check the robot's power status and try restarting it. Is there anything else I can help you with?"
[0446] 4. Examples of prompts
[0447] "My robot has stopped working, what should I do?"
[0448] "How do I install the new parts?"
[0449] This allows the user to receive accurate work guidance in real time, improving the efficiency and safety of robot operation, and serves as a clear solution to the problem that the invention aims to solve.
[0450] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0451] Step 1:
[0452] The server hosts the chatbot application and loads the generative AI model and emotion recognition engine. The device installs and launches the chatbot application. The user interface is displayed, allowing the user to enter a message. When the user launches the application, the initial screen displays, "Hello. Is there anything I can help you with? How can I help you operate your robot?"
[0453] Step 2:
[0454] The user logs in to the application. The server checks the user information and retrieves the user profile. The retrieved user profile includes past operation history and setting information. The device displays a personalized UI based on the user's settings and past operation history. This allows the device to provide the user with appropriate suggestions and guidance.
[0455] Step 3:
[0456] The user enters questions or instructions into the chat box. For example, "How do I install a new part?" The system also analyzes the user's voice and facial expressions using an emotion engine. The device then sends these inputs to the server. The entered text data and voice data are then sent to the server.
[0457] Step 4:
[0458] The server uses a generative AI model to analyze natural language and an emotion recognition engine to recognize the user's emotional state. The analysis results identify the user's intentions, requests, and emotional state. For example, it may determine that the user wants to know the procedure for installing a new part. The user is feeling a little anxious. Natural language analysis extracts keywords from the input text data, and emotion recognition is performed by analyzing voice tone and facial expression data.
[0459] Step 5:
[0460] The server generates an appropriate response based on the analysis results and the user's emotional state. For example, it may describe the installation procedure in detail in a calm tone. It may generate a response such as, "The installation procedure is as follows: First, turn off the power and check the safety devices. Then, install the parts. Please let us know if you have any questions." The generated response is in text format.
[0461] Step 6:
[0462] The generated response is sent from the server to the terminal. The terminal displays the received response on the user interface. Based on the displayed message, the user can obtain specific instructions for performing the next action. For example, a message such as "The installation procedure is as follows. First, turn off the power and check the safety devices. Then install the parts. Please let us know if you have any questions" may be displayed.
[0463] Step 7:
[0464] When a user reports a specific problem, for example, "My robot has stopped working, what should I do?", the server analyzes the problem and evaluates the user's emotional state using an emotion recognition engine. The generated troubleshooting guidance is sent to the device as specific solution steps. The device displays the solution steps on the user interface and provides emotion-sensitive instructions. For example, a message such as "Please check the robot's power status and try restarting it. Is there anything else I can help you with?" is displayed.
[0465] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0466] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0467] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0468] [Second embodiment]
[0469] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0470] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0471] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0472] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0473] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0474] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0475] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0476] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0477] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0478] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0479] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0480] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0481] As an embodiment of the present invention, we provide a chatbot system that supports the use of operating systems (OS) for PCs and smartphones. The specific operation method and processing content of this system are described below.
[0482] First, the system includes means for receiving natural language input from a user, means for analyzing the received input, means for generating a response based on the analysis result, and means for displaying the response to the user.
[0483] Program processing overview
[0484] 1. Initialization and Setup
[0485] The server hosts the chatbot application, loads the generative AI model, and exposes API endpoints for interacting with the OS and users, starting the chatbot.
[0486] The device launches a dedicated chatbot app built into the OS, which displays a user interface (UI) and allows the user to enter a message.
[0487] The user accesses the chatbot application and starts a dialogue. The chatbot receives the natural language input from the user in text format.
[0488] 2. User Authentication and Profiling
[0489] A user logs in to an application. The server checks the user information and retrieves the user profile. The retrieved user profile includes past operation history and setting information. This allows personalized suggestions to be made to the user.
[0490] The device will prepare individually customized suggestions and guides based on the user's settings and past operation history.
[0491] 3. Natural Language Understanding and Dialogue Generation
[0492] Users enter questions or instructions into a chat box (e.g., "I want to install a new app"). The device sends this input to the server, which uses generative AI models to parse the natural language and generate an appropriate response.
[0493] The generated response is sent to the terminal and displayed on the user interface, allowing the user to intuitively understand the procedures for operating the OS.
[0494] 4. Function execution and operation support
[0495] Based on the response from the server, the device will take specific action. For example, if the user inputs "I want to install a new app," the server will search for the appropriate application and generate a response guiding the user through the installation process. The device will then display this response and guide them through the necessary operations.
[0496] The user follows the instructions to perform the action (e.g., clicks the install button). The device tracks the operation status in real time and displays additional guidance as needed.
[0497] 5. Troubleshooting
[0498] When a user reports a specific problem (e.g., "I can't connect to Wi-Fi"), the server analyzes the problem and identifies possible causes. The generated troubleshooting guidance is sent to the device as specific solution steps. The device displays the solution steps on the user interface and provides detailed instructions to the user.
[0499] Specific examples
[0500] New user onboarding
[0501] User: Starts up a newly purchased PC and uses the chatbot for the first time.
[0502] Device: Display "Nice to meet you. Let's go over the basic settings together. First, let's set the language."
[0503] User: Type "Please set to Japanese."
[0504] Terminal: Send this input to the server.
[0505] Server: Analyzes the message and generates instructions to change the language setting to Japanese.
[0506] Device: Change the language setting to Japanese and notify the user that "Settings are complete."
[0507] Introducing and explaining new features
[0508] User: After an update, launch the chatbot and type, "Tell me about the new features."
[0509] Server: Searches for information about new features and generates a summary of them.
[0510] On your device: "The new update includes a refreshed design, battery optimizations, and new security features."
[0511] Fault diagnosis and troubleshooting
[0512] User: Reports "No sound."
[0513] Server: Analyzes the problem, lists possible causes, and generates solutions.
[0514] Device: "Possible causes of no sound include the volume setting, muting, or speaker failure. First, check the volume setting." is displayed.
[0515] Users: Follow the instructions and double-check your volume settings.
[0516] In this way, a dedicated chatbot can assist users in various situations, making operating the OS more intuitive and efficient. This system allows users to effectively use the OS's functions and settings, and quickly respond to any problems that arise.
[0517] The processing flow will be explained below.
[0518] Step 1:
[0519] The server launches the chatbot application and loads the generative AI model, which is a machine learning model for natural language analysis and response generation. The server also opens an API endpoint and prepares to accept connection requests from devices.
[0520] Step 2:
[0521] The device installs and launches the chatbot application. The user interface (UI) is displayed, allowing the user to enter a message. The initial screen displays, "Nice to meet you. Welcome to the chatbot. How can I help you?"
[0522] Step 3:
[0523] The user enters login information (e.g., user ID and password) into the UI and clicks the "Login" button, which sends a login request from the terminal to the server.
[0524] Step 4:
[0525] The server compares the received login information with a database and performs authentication. If authentication is successful, the server obtains the user profile and sends it to the terminal. The user profile includes past operation history and individual setting information.
[0526] Step 5:
[0527] The device will then display a personalized UI based on the received user profile, such as shortcuts to frequently used functions and recently accessed applications.
[0528] Step 6:
[0529] The user types a question or instruction into the chat box (e.g., "I want to install a new app"), and the message is sent directly to the server.
[0530] Step 7:
[0531] The server receives the user's input message and uses a generative AI model to perform natural language analysis, identifying the user's intent and request (e.g., "I want to know the steps to install a new app").
[0532] Step 8:
[0533] The server generates an appropriate response based on the analysis results, such as a command to open the application store or guidance on installation procedures.
[0534] Step 9:
[0535] The generated response is sent from the server to the device, and may include specific instructions or options.
[0536] Step 10:
[0537] The device displays the received response on the user interface, for example, a confirmation message such as "Do you want to open the application store?"
[0538] Step 11:
[0539] The user follows the displayed instructions (e.g., answers "yes" and opens the application store), and the user's selection is sent back to the server.
[0540] Step 12:
[0541] Based on the user's selection, the server determines the next action and sends the necessary instructions to the terminal, such as to display search results for a particular application.
[0542] Step 13:
[0543] The device updates its user interface based on the instructions received from the server, for example by displaying a list of search results and allowing the user to select the app they want to install.
[0544] Step 14:
[0545] The user selects the app they want to install and sends an installation request to the server again.
[0546] Step 15:
[0547] The server confirms the installation request, initiates the necessary installation process, and notifies the device of the installation progress.
[0548] Step 16:
[0549] The terminal displays the installation progress to the user and notifies the user when the installation is complete.
[0550] Step 17:
[0551] The user launches the application after the installation is complete and begins using it.
[0552] Through this process, users can interact with the OS in natural language and smoothly perform the required actions.
[0553] Example 1
[0554] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0555] In recent years, the use of information devices such as PCs and smartphones has continued to expand, resulting in a demand for support regarding device operation and settings. However, conventional support systems have difficulty providing personalized support based on individual users' needs and operation history, limiting the improvement of usability. Furthermore, when it comes to troubleshooting, it is difficult to provide a prompt and accurate response to specific problems users face.
[0556] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0557] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, means for generating a corresponding response based on the analysis, means for displaying the response to the user, means for opening an API endpoint for communication with an OS and a user, means for loading a generative AI model, means for analyzing the user's input and generating an appropriate response, means for sending the generated response to a terminal, means for customizing a user interface, means for tracking user operations in real time and displaying additional guidance, and means for analyzing problems, identifying possible causes, and generating solutions. This enables personalized support based on the user's individual needs and operation history, and also enables quick and accurate troubleshooting.
[0558] "Natural language input from a user" refers to a message expressed in everyday language that is input by a user in the form of text, voice, or the like.
[0559] "Means for analyzing natural language input" refers to an algorithm or system that understands the natural language message received from the user and analyzes its grammatical structure and meaning.
[0560] "Means for generating a corresponding response" refers to a system or algorithm for generating an appropriate reply or instruction based on the results of natural language analysis.
[0561] "Means for displaying a response to a user" refers to an interface or device for displaying the generated response so that the user can visually confirm it.
[0562] "Means for opening API endpoints" refers to a system that provides the protocols and interfaces necessary to communicate with information terminals and operating systems.
[0563] "Means for loading a generative AI model" refers to the process or algorithm used to load a generative model into memory or a processor and make it executable.
[0564] "Means for analyzing user input and generating an appropriate response" refers to a process or system that analyzes user input and generates an optimal response based on the results.
[0565] The "means for transmitting the generated response to the terminal" refers to a system for sending the generated response data to the terminal via a network.
[0566] "Means for customizing a user interface" refers to a method or system for changing the design and functionality of an interface based on user settings and operation history.
[0567] "Means for tracking user operations in real time and displaying additional guidance" refers to a system that monitors the operations performed by a user in real time and provides next steps or additional instructions as needed.
[0568] "Means for analyzing problems, identifying possible causes, and generating solutions" refers to the systems and processes used to analyze user-reported problems, identify their causes, and generate appropriate solutions.
[0569] This invention relates to a chatbot system that supports the use of operating systems (OS) for PCs and smartphones. Specifically, it is a system that uses a generative AI model to analyze natural language input from users and generate appropriate responses.
[0570] This system mainly consists of three elements: a server, a terminal, and a user.
[0571] server
[0572] The server hosts the chatbot application and loads the generative AI model (e.g., GPT-4). The server exposes API endpoints for communication with the OS and users, and receives and analyzes user input. The server has the following capabilities:
[0573] A means of receiving natural language input from the user
[0574] A means of parsing natural language input
[0575] A means of generating a corresponding response based on the analysis results
[0576] A means of sending the generated response to the terminal
[0577] A means of analyzing problems, identifying possible causes, and generating solutions
[0578] Specifically, the server performs the following process.
[0579] 1. When receiving input from the user, retrieve the data through the appropriate API endpoint.
[0580] 2. The received input data is analyzed using a generative AI model, which performs tokenization and grammar analysis to understand meaning.
[0581] 3. Based on the analysis results, the server generates an appropriate response, referencing the user profile and past operation history to provide personalized suggestions.
[0582] 4. The generated response is sent to the device through the API endpoint.
[0583] Terminal
[0584] The device runs a dedicated chatbot application built into the operating system. The device is equipped with the following means:
[0585] A way to customize the user interface
[0586] A means to track user actions in real time and display additional guidance
[0587] Specifically, the terminal operates as follows.
[0588] 1. Launch the chatbot application and display a user interface that allows the user to enter a message.
[0589] 2. If there is any input from the user, send it to the server.
[0590] 3. Display the response received from the server in the user interface, providing interactive instructions as needed.
[0591] 4. Track user actions in real time and provide appropriate guidance when information is missing or additional guidance is needed.
[0592] User
[0593] Users access the chatbot application using a computer or smartphone and receive support through natural language dialogue. Users mainly perform the following operations:
[0594] 1. Log in to the application and provide the required authentication information.
[0595] 2. Enter questions or instructions in the chat box and receive responses from the server.
[0596] 3. Follow the displayed instructions and guides to operate the OS.
[0597] Specific examples
[0598] For example, when a user who has just purchased a new PC performs initial setup, the following dialogue may take place:
[0599] User: "I just started up my new PC. What do I do now?"
[0600] Server: "Nice to meet you. Let's go over some basic setup together. First, let's set up the language."
[0601] User: "Please set it to Japanese."
[0602] Server: "The language setting has been changed to Japanese. Next, set the time zone."
[0603] Also, if the user wants to install a new application, the prompt sentence is as follows:
[0604] Prompt for "I want to install a new app":
[0605] Input to the server: "The user wants to install a new app. Please guide them through the installation process."
[0606] Example response: "Find the right application and guide me through the installation process."
[0607] This invention allows users to intuitively and efficiently use OS functions and settings, and also allows them to receive prompt and accurate support when problems occur.
[0608] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0609] Step 1: Initialization and Setup
[0610] server:
[0611] The server hosts the chatbot application, loads the generative AI model (e.g., GPT-4), opens an API endpoint, and prepares to communicate with the OS and the user. Specific operations include the following steps:
[0612] The server reads specific configuration files and database connection information.
[0613] The server loads the generative AI model into memory and makes it executable.
[0614] Device:
[0615] The device launches a dedicated chatbot app and displays a user interface (UI), allowing the user to enter a message. Specifically, the process involves the following steps:
[0616] The device prepares the UI for display, rendering text boxes, submit buttons, etc.
[0617] The terminal checks the connection to the server and whether communication is possible.
[0618] User:
[0619] A user accesses a chatbot application and starts a conversation. The specific operations include the following steps:
[0620] The user checks the application's UI and enters an initial message such as "Hello."
[0621] Step 2: User authentication and profiling
[0622] User:
[0623] A user logs into an application, typically by entering a username and password, and sometimes by using biometric authentication such as facial recognition or fingerprint authentication.
[0624] server:
[0625] The server checks the user's authentication information against the database and retrieves the user profile. Specifically, the following process is performed:
[0626] The server takes the entered username and password and issues an SQL query to the database to verify it.
[0627] If the authentication is successful, the server obtains the user profile (including operation history and setting information) and stores it in memory.
[0628] Device:
[0629] The device customizes the UI based on the user profile information received from the server. Specifically, the following process is performed:
[0630] The device changes the UI style based on the user's language and theme settings.
[0631] Step 3: Natural Language Understanding and Dialogue Generation
[0632] User:
[0633] The user types a question or instruction into the chat box, for example, "I want to install a new app."
[0634] Device:
[0635] The terminal sends the user's input to the server. The specific procedure is as follows:
[0636] The device receives the input as text data and sends an HTTP POST request through the API endpoint.
[0637] server:
[0638] The server analyzes the received user input using a generative AI model. Specifically, the following data processing and calculations are performed:
[0639] The server tokenizes the input data, performs grammatical analysis, and analyzes it to understand its meaning.
[0640] Based on the analysis results, an appropriate response is generated.
[0641] Device:
[0642] The generated response is sent to the terminal and displayed on the user interface. Specifically, the process includes the following steps:
[0643] The terminal acquires the response data received from the server and displays it in the text box.
[0644] Step 4: Implementing functions and assisting operations
[0645] server:
[0646] The server then generates a response that recommends a specific action, such as instructions for installing a new app.
[0647] Device:
[0648] The terminal guides the user according to the procedure from the server. Specifically, the terminal performs the following procedure.
[0649] The device will display instructions such as "Click the install button."
[0650] User:
[0651] The user follows the instructions displayed on the device, for example, "click the install button."
[0652] Step 5: Troubleshooting
[0653] User:
[0654] The user reports a specific problem, for example, "I can't connect to Wi-Fi."
[0655] server:
[0656] The server analyzes the problem, identifies possible causes, and generates solutions. Specifically, the following data processing and calculations are performed:
[0657] The server references the data for troubleshooting and identifies the cause through natural language analysis.
[0658] The server generates solutions and organizes them into guidance.
[0659] Device:
[0660] The generated troubleshooting guidance is displayed in the user interface, specifically by following the steps below.
[0661] The device will display instructions such as "First, please check your volume settings."
[0662] This system allows users to operate the OS and solve problems intuitively and efficiently.
[0663] (Application example 1)
[0664] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0665] Conventional systems make it difficult for users to operate and configure the system in natural language, and they lack sufficient support for security-related settings and troubleshooting. Furthermore, they do not provide personalized suggestions that effectively utilize user profiles and operation histories, making it difficult for users to operate the system intuitively and efficiently.
[0666] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0667] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, and means for generating a corresponding response based on the analysis, thereby enabling the user to perform security settings and other operations in natural language.
[0668] Furthermore, the server includes means for performing specific actions based on the generated responses and means for analyzing problems and generating solutions during troubleshooting, thereby enabling the user to obtain specific troubleshooting and operation guidance.
[0669] Furthermore, by including a means for acquiring and managing a user profile, a means for providing personalized suggestions to the user based on the profile, and a means for adjusting security settings based on the user's operation history and setting information, intuitive and individually optimized operation becomes possible, allowing the user to efficiently manage security settings.
[0670] "Natural language input" is an input format in which a user gives instructions or asks questions to a system in everyday language.
[0671] "Analysis" refers to the process of converting input natural language into a form that a machine can understand and interpreting its meaning.
[0672] A "response" is a reply or instruction from the system that is generated based on the parsed natural language input.
[0673] A "user profile" is a data set that collects information about a user, including individual settings and past operation history.
[0674] "Personalized suggestions" are suggestions that individually provide users with information and guidance that is most suitable for them based on their user profile.
[0675] The "operation history" is a record of a series of operations and setting changes that the user has performed in the past.
[0676] A "specific action" is a specific operation or process that the system actually performs based on a user's instruction.
[0677] "Troubleshooting" is the process of identifying the cause of a problem and providing a solution.
[0678] A "generative AI model" is an artificial intelligence technology that uses machine learning to understand natural language and generate appropriate responses.
[0679] "Security settings" refers to a set of settings and policies that ensure the security of a device or system.
[0680] As an embodiment of the present invention, we will explain a chatbot system in which a server and a terminal cooperate to assist users in security configuration and troubleshooting. This system utilizes a generative AI model to provide appropriate responses to security-related instructions and questions entered by the user in natural language.
[0681] Programs and hardware / software used
[0682] An API (specifically, the OpenAI API) that runs the generative AI model is installed on the server, which analyzes input from users and generates responses. A smartphone is used as the terminal, and a chatbot app is installed on this smartphone. Users interact with the system through this app.
[0683] Natural language input processing begins when a user enters instructions or questions in natural language into the chatbot app. This input is sent to the server as text data, where it is analyzed by the generative AI model. Based on the results of the analysis, the server generates an appropriate response and sends it back to the device.
[0684] User profile management is achieved by applications sending user information to a server, which then stores the user's operation history and settings information and generates personalized suggestions and actions based on this information, which can be useful when adjusting security settings.
[0685] For example, if a user types "check Wi-Fi settings," the device will execute a script based on the generated response to help check the actual Wi-Fi settings, allowing users to easily troubleshoot security settings and resolve issues.
[0686] When troubleshooting, users report their issues in natural language and their input is sent to the server, which uses generative AI models to analyze the cause of the problem and generate a solution as a response. For example, if a user reports "no sound," specific steps such as "check your volume settings" are provided.
[0687] Specific examples
[0688] User: "Check the security status of your device"
[0689] Chatbot: "We're running a security check on your device. Please wait a moment... All security settings are OK."
[0690] Example prompt sentence:
[0691] User input: Check the security status of your device
[0692] Security Chatbot Response: We're going to perform a security check on your device. Please wait... All your security settings are fine.
[0693] In this way, users can easily perform security-related operations in natural language through the chatbot and receive appropriate troubleshooting guidance when problems arise, enabling users to manage their security settings intuitively and efficiently.
[0694] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0695] Step 1:
[0696] The device launches the chatbot app, and the user inputs instructions or questions in natural language. The input text data is sent to the server via an API call. The server receives the text data and prepares it for analysis by the generative AI model.
[0697] Input: User's natural language input (text data)
[0698] Output: Text data sent to the server
[0699] Step 2:
[0700] The server passes the received natural language input to the generative AI model, which begins the analysis process. The generative AI model analyzes the input natural language, understands its meaning, and generates an appropriate response. The analysis uses the OpenAI API.
[0701] Input: Text data sent to the server
[0702] Output: Analysis results from the generative AI model
[0703] Step 3:
[0704] Based on the generated analysis, the server generates a specific response to the user's instruction, such as "check your Wi-Fi settings" or any other troubleshooting steps required.
[0705] Input: Analysis results from generative AI model
[0706] Output: Response data to the user
[0707] Step 4:
[0708] The server sends the generated response data to the terminal, which receives the response and displays it on the user interface within the chatbot app. The user can then view it and take the next action.
[0709] Input: Response data to the user
[0710] Output: Response data sent to the terminal
[0711] Step 5:
[0712] The user takes action based on the generated response. For example, if the instruction is "Check your Wi-Fi settings," the user opens the Wi-Fi settings screen and adjusts the settings. After completing this action, the user again provides feedback to the chatbot.
[0713] Input: Response data displayed to the user
[0714] Output: User's actual actions
[0715] Step 6:
[0716] If the user wants to input additional instructions to resolve the issue, the process repeats from step 1. For example, if the user inputs "Please also check the volume settings," the server will analyze it again and return the appropriate instructions.
[0717] Input: Additional user instructions (natural language input)
[0718] Output: The analysis process is restarted.
[0719] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0720] As an embodiment of the present invention, we provide a chatbot system that supports the use of operating systems (OS) for PCs and smartphones. In particular, we will explain an embodiment that incorporates an engine that recognizes the user's emotions.
[0721] First, the system has the means to receive natural language input from users, parse it, generate and display a response, and, combined with an emotion engine, recognizes the user's emotional state and adapts responses based on that, achieving a high level of personalization.
[0722] Program processing overview
[0723] 1. Initialization and Setup
[0724] The server hosts the chatbot application and loads the generative AI model, a machine learning model for natural language analysis and response generation, and the emotion engine, a model for recognizing user emotions.
[0725] The chatbot application is installed on the device and launched. The user interface (UI) is displayed, allowing the user to enter a message. The initial screen displays, "Hello. Is there anything I can help you with?"
[0726] The user accesses the chatbot application and starts a dialogue, receiving the user's natural language or voice input in text format.
[0727] 2. User Authentication and Profiling
[0728] A user logs in to an application. The server checks the user information and retrieves the user profile. The retrieved user profile includes past operation history and setting information.
[0729] The device displays a personalized UI based on the user's settings and past operation history, allowing it to provide more appropriate suggestions and guidance to the user.
[0730] 3. Natural Language Understanding and Emotion Recognition
[0731] Users enter questions or instructions into the chat box (e.g., "I want to install a new app"). The system also analyzes the user's facial expressions and tone of voice using an emotion engine. The device then sends these inputs to the server.
[0732] The server uses a generative AI model to analyze natural language and an emotion engine to recognize the user's emotions. The analysis results identify the user's intent, request, and emotional state (e.g., "I want to know the steps to install a new app. The user is feeling a bit anxious and worried.").
[0733] 4. Response generation and emotional adaptation
[0734] Based on the analysis and emotional state, the server generates an appropriate response, such as a command to open the application store or friendly, gentle installation instructions.
[0735] The generated response is sent from the server to the device, and may include specific instructions or an emotionally sensitive tone.
[0736] The device displays the received response on the user interface, for example, a message such as "Open the application store and install the desired app. If you have any questions, please feel free to ask."
[0737] 5. Troubleshooting
[0738] When a user reports a specific problem (e.g., "I can't connect to Wi-Fi"), the server analyzes the problem and evaluates the user's emotional state using an emotion engine. The generated troubleshooting guidance is sent to the device as specific solution steps. The device then displays the solution steps on the user interface and provides emotion-sensitive instructions.
[0739] Specific examples
[0740] New user onboarding
[0741] User: Starts up a newly purchased PC and uses the chatbot for the first time.
[0742] Device: Display "Hello. Let's go through the basic setup together. If you have any questions, please feel free to ask."
[0743] User: Type "Please set to Japanese."
[0744] Terminal: Send this input to the server.
[0745] Server: Parses the message and generates instructions to change the language setting to Japanese. Recognizes that the user is feeling a bit anxious and generates a response in a friendly tone.
[0746] Device: Change the language setting to Japanese and notify the user, "Setup complete. You can now use Japanese. Is there anything else we can help you with?"
[0747] Introducing and explaining new features
[0748] User: After an update, launch the chatbot and type, "Tell me about the new features."
[0749] Server: Searches for information about new features and generates summaries of them, recognizing the user's emotional state and crafting interesting descriptions.
[0750] On your device: "Our new update features a fresh design, battery optimizations, and new security features."
[0751] Fault diagnosis and troubleshooting
[0752] User: Reports "No sound."
[0753] Server: Analyzes the problem and recognizes the user's emotional state using an emotion engine. Lists possible causes and generates solutions.
[0754] Device: "Possible reasons for the lack of sound include the volume setting, muting, or a faulty speaker. First, check the volume setting. If you have any questions, please let us know."
[0755] User: Follow the instructions and check the volume settings again. In this way, a system incorporating an emotion engine allows users to receive responses that take their emotions into consideration, allowing them to operate the OS with greater peace of mind.
[0756] The processing flow will be explained below.
[0757] Step 1:
[0758] The server launches the chatbot application and loads the generative AI model and emotion engine. The generative AI model is a machine learning model for natural language analysis and response generation, and the emotion engine is a model for recognizing user emotions. The server also opens an API endpoint and prepares to accept connection requests from devices.
[0759] Step 2:
[0760] The chatbot application is installed on the device and launched. The user interface (UI) is displayed, allowing the user to enter a message. The initial screen displays, "Hello. Is there anything I can help you with?"
[0761] Step 3:
[0762] The user enters login information (e.g., user ID and password) into the UI and clicks the "Login" button, which sends a login request from the terminal to the server.
[0763] Step 4:
[0764] The server compares the received login information with a database and performs authentication. If authentication is successful, the server obtains the user profile and sends it to the terminal. The user profile includes past operation history and individual setting information.
[0765] Step 5:
[0766] The device will then display a personalized UI based on the received user profile, such as shortcuts to frequently used functions and recently accessed applications.
[0767] Step 6:
[0768] Users enter questions or instructions into the chat box (e.g., "I want to install a new app"). The entered message is sent directly to the server. The system also analyzes the user's facial expressions and tone of voice using an emotion engine.
[0769] Step 7:
[0770] The server receives the user's input message and uses a generative AI model to perform natural language analysis, identifying the user's intent and request (e.g., "I want to know the steps to install a new app").
[0771] Step 8:
[0772] The server uses an emotion engine to recognize the user's emotional state, for example, to determine whether the user is feeling anxious or worried.
[0773] Step 9:
[0774] The server generates an appropriate response based on the analysis and emotional state, for example, "Do you want to open the app store?" in a calm and reassuring tone.
[0775] Step 10:
[0776] The generated response is sent from the server to the device, and may include specific instructions or an emotionally sensitive tone.
[0777] Step 11:
[0778] The device displays the received response on the user interface, for example, "Open the application store and install the desired app. If you have any questions, please feel free to ask."
[0779] Step 12:
[0780] The user follows the displayed instructions (e.g., answers "yes" and opens the application store), and the user's selection is sent back to the server.
[0781] Step 13:
[0782] Based on the user's selection, the server determines the next action and sends the necessary instructions to the terminal, such as to display search results for a particular application.
[0783] Step 14:
[0784] The device updates its user interface based on the instructions received from the server, for example by displaying a list of search results and allowing the user to select the app they want to install.
[0785] Step 15:
[0786] The user selects the app they want to install and sends an installation request to the server again.
[0787] Step 16:
[0788] The server confirms the installation request, initiates the necessary installation process, and notifies the device of the installation progress.
[0789] Step 17:
[0790] The terminal displays the installation progress to the user and notifies the user when the installation is complete.
[0791] Step 18:
[0792] Once the installation is complete, the user launches the application and begins using it. Through this process, the user can interact with the OS based on natural language and emotions to smoothly perform the required actions.
[0793] Example 2
[0794] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0795] Conventional chatbot systems have difficulty generating appropriate responses to users' natural language input, and in particular, they are unable to recognize emotions and reflect them in responses, which means they are unable to provide a satisfying conversational experience for users.In addition, they lack personalized suggestions that utilize the user's profile information and past operation history, making it difficult to achieve user-friendly operation.
[0796] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0797] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, means for generating a corresponding response based on the analysis, means for displaying the response to the user, means for analyzing natural language using a generative AI model, means for identifying the user's emotional state using an emotion recognition engine, means for generating an appropriate response based on the analysis result and the emotional state, and means for displaying the response in a tone that takes the user's emotions into consideration. This makes it possible to provide more personalized suggestions and guidance while taking the user's emotions into consideration, thereby achieving a highly satisfying dialogue experience.
[0798] "Natural language input" refers to text or voice input entered by a user in the form of everyday spoken language.
[0799] A "generative AI model" refers to an algorithmic model that uses machine learning technology to analyze natural language and generate responses.
[0800] "Emotion recognition engine" refers to a software module for analyzing and identifying an emotional state from a user's input.
[0801] A "user profile" refers to a collection of data including a user's personal information, past operation history, and setting information.
[0802] "Personalization" refers to providing customized suggestions and responses based on individual user characteristics and past behavior.
[0803] "Operation history" refers to a record of operations that a user has performed in the past.
[0804] "Setting information" refers to information about various settings that a user makes when using an application or system.
[0805] A "server" refers to a computer or software system used to receive requests from users and process or respond to data.
[0806] "Terminal" refers to the device that the user actually operates (e.g., a PC or smartphone).
[0807] "User interface" refers to the screen display and input means that allow a user to interact with a system.
[0808] "Login" refers to the authentication procedure required for a user to access a system.
[0809] "Analysis" refers to the processing of data to understand a user's natural language input and emotional state.
[0810] A "response" refers to a reply or instruction that a system generates and provides in response to a user's input.
[0811] As an embodiment of the present invention, we provide a chatbot system that supports the use of operating systems (OS) for personal computers and smartphones. This system incorporates a generative AI model and an emotion recognition engine.
[0812] First, the server hosts the chatbot application and loads a generative AI model (e.g., GPT-3.5) and an emotion recognition engine (e.g., a popular emotion recognition API). The generative AI model is used to parse natural language and generate responses, and the emotion recognition engine is used to recognize the user's emotional state. These models and engines are initialized by the server by loading the API key and model data.
[0813] The device then allows the user to install and launch the chatbot application. When the user launches the application, a user interface appears, displaying a welcome message that reads, "Hello, how can I help you?"
[0814] The user accesses the chatbot application and starts a dialogue using natural language or voice. The input content is received in text format and sent from the terminal to the server. If the user inputs voice, the voice is converted into text and sent to the server.
[0815] When a user logs in to an application, the server checks the user information and obtains profile data (past operation history, setting information, etc.). Based on the obtained profile data, a personalized UI is displayed to the user. For example, frequently used functions can be displayed as shortcuts based on the past operation history.
[0816] The server analyzes the natural language input from the user using a generative AI model to identify the user's intentions and requests. It also uses an emotion recognition engine to recognize the user's emotional state. For example, if a user inputs "I want to install a new app," the analysis results might determine that "I want to know the steps to install a new app, and I'm feeling a bit impatient and anxious."
[0817] Based on the analysis results and the emotional state, the server uses a generative AI model to generate an appropriate response, such as "Please visit the application store and install the app you want. If you have any questions, please feel free to ask." This response is generated in an emotionally sensitive tone.
[0818] The generated response is sent to the device and displayed in the user interface. The displayed response is provided in an emotionally sensitive tone, allowing the user to proceed with the operation with confidence.
[0819] When a user reports a specific problem, for example, "I can't connect to Wi-Fi," the server analyzes the problem and uses an emotion recognition engine to assess the user's emotional state. Based on the analysis, specific solutions are generated and sent to the device. For example, a message might read, "Possible causes of the lack of audio include volume settings, muting, or a faulty speaker. First, check your volume settings."
[0820] Specific examples
[0821] New user onboarding
[0822] User: Starts up a newly purchased PC and uses the chatbot for the first time.
[0823] Device: Display "Hello. Let's go through the basic setup together. If you have any questions, please feel free to ask."
[0824] User: Type "Please set to Japanese."
[0825] Terminal: Send this input to the server.
[0826] Server: Parses the message and generates instructions to change the language setting to Japanese. Recognizes that the user is feeling a bit anxious and generates a response in a friendly tone.
[0827] Device: Change the language setting to Japanese and notify the user, "Setup complete. You can now use Japanese. Is there anything else we can help you with?"
[0828] Introducing and explaining new features
[0829] User: After an update, launch the chatbot and type, "Tell me about the new features."
[0830] Server: Searches for information about new features and generates summaries of them, recognizing the user's emotional state and crafting interesting descriptions.
[0831] On your device: "Our new update features a fresh design, battery optimizations, and new security features."
[0832] Fault diagnosis and troubleshooting
[0833] User: Reports "No sound."
[0834] Server: Analyzes the problem and recognizes the user's emotional state using an emotion recognition engine. Lists possible causes and generates solutions.
[0835] Device: "Possible reasons for the lack of sound include the volume setting, muting, or a faulty speaker. First, check the volume setting. If you have any questions, please let us know."
[0836] In this way, a system incorporating an emotion recognition engine allows users to receive responses that take their emotions into consideration, allowing them to operate the OS with greater peace of mind.
[0837] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0838] Step 1:
[0839] The server hosts the chatbot application and loads the generative AI model (e.g., GPT-3.5) and emotion recognition engine (e.g., a popular emotion recognition API). Specifically, it loads the API key and model data to initialize the model. The input at this stage is the required API key and model data, and the output is a ready-to-use generative AI model and emotion recognition engine.
[0840] Step 2:
[0841] When a user installs and launches a chatbot application, the device loads the user interface. The initial screen displays "Hello, is there anything I can help you with?" When a user launches the application, the input is the application launch command, and the output is a user interface displaying the message.
[0842] Step 3:
[0843] The user inputs questions or instructions to the chatbot using natural language or voice. For example, they might input "I want to install a new app." This input is received by the device in text format. If the input is voice, the voice is converted to text. At this stage, the input is the user's natural language or voice input, and the output is text data.
[0844] Step 4:
[0845] The device sends the received text data to the server. It also sends text data converted from speech. Specifically, it sends the text data to the server via an HTTP request. At this stage, the input is user data in text format, and the output is the data sent to the server.
[0846] Step 5:
[0847] The server analyzes the received text data using a generative AI model to identify the user's intentions and requests. It also uses an emotion recognition engine to recognize the user's emotional state. For example, if a user inputs "I want to install a new app," the analysis results might determine that "I want to know the steps to install a new app. The user is feeling a bit impatient and anxious." The input at this stage is text data, and the output is the analysis results regarding the user's intentions and emotional state.
[0848] Step 6:
[0849] The server uses a generative AI model to generate an appropriate response based on the analysis results and the emotional state. For example, it generates a message such as, "Open the application store and install the app you want. If you have any questions, please feel free to ask." This response is generated in an emotionally sensitive tone. The input at this stage is the analysis results and the emotional state data, and the output is an appropriate response message.
[0850] Step 7:
[0851] The server sends the generated response message to the device. The device receives the response and displays it on its user interface. The displayed message contains specific instructions such as "Open the application store and install the desired app." The input at this stage is the response message, and the output is the message displayed on the user interface.
[0852] Step 8:
[0853] When a user reports a specific problem, for example, "I can't connect to Wi-Fi," the server analyzes the problem and uses an emotion recognition engine to assess the user's emotional state. The analysis results in a list of possible causes and generates a solution. For example, a message might read, "Possible causes of the no sound include the volume setting, muting, or a faulty speaker. Please check the volume setting first." The input at this stage is text data about the specific problem, and the output is the problem analysis and a solution message.
[0854] Step 9:
[0855] The terminal displays the solution message received from the server on the user interface. For example, it displays "Please check your volume settings." This message is expressed in an emotionally sensitive tone. The input at this stage is the solution message, and the output is the solution message displayed on the user interface.
[0856] This allows users to receive emotionally sensitive responses and operate with peace of mind.The system improves the user experience by combining natural language input analysis and emotion recognition.
[0857] (Application example 2)
[0858] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0859] In factory work environments, workers are required to be able to operate robots and troubleshoot them quickly and reliably. However, conventional systems have difficulty responding to the user's emotional state and real-time work environment, which often causes stress and confusion for workers. This can hinder efficient work execution and reduce safety.
[0860] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0861] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, means for generating a corresponding response, means for recognizing the user's emotional state, and means for adapting the response based on the user's emotional state. This enables generation of a response that takes the user's emotions into consideration. The server also includes means for acquiring the user's operation history and setting information, means for providing information related to the work environment in real time, and means for providing appropriate operation guidance to the user. This enables appropriate guidance based on real-time information about the work environment, allowing work to be performed safely and efficiently.
[0862] "Natural language input" refers to a language that is input by a user in the form of text, speech, or other language that is commonly used by humans.
[0863] "Analysis" refers to the syntactic and semantic processing and interpretation of input natural language.
[0864] A "response" is a response to a question or instruction that is generated based on the results of the analysis.
[0865] "Display" refers to providing the generated response to the user visually or audibly.
[0866] "Emotional state" refers to the user's psychological and emotional state as determined by their tone of voice and facial expression.
[0867] A "profile" is a collection of data about a user, including the user's personal information, operation history, setting information, and the like.
[0868] "Personalization" refers to providing individualized offers and information based on a user's profile information.
[0869] An "operation history" is a record of operations and actions that a user has performed in the past.
[0870] "Setting information" refers to setting items and parameters that are individually set by the user for the system or application.
[0871] "Work environment information" is information relating to the user's real-time work location and situation.
[0872] "Operational guidance" refers to procedures or instructions provided to a user to perform a specific task or operation.
[0873] This embodiment of the present invention is a chatbot system with advanced response capabilities incorporating user emotion recognition. It is particularly intended for supporting robotic operation in factory work environments. The system analyzes the user's natural language input, generates corresponding responses, and displays those responses. It also provides appropriate operational guidance in real time based on the user profile and work environment information.
[0874] 1. Hardware and Software Used
[0875] This system uses the following hardware and software:
[0876] Hardware: Smartphones, smart glasses, head-mounted displays
[0877] software:
[0878] Python: To execute the main logic of the program
[0879] OpenAI API: Natural Language Analysis and Response Generation
[0880] Emotion Recognition API: To recognize user emotions
[0881] 2. Program processing content
[0882] The server uses a generative AI model to analyze natural language and generate appropriate responses to user input. Furthermore, it uses the Emotion Recognition API to determine the user's emotional state and adapts responses based on that emotion. Specific responses include:
[0883] 3. System Operation Example
[0884] Operation Guidance
[0885] When a user inquires about how to operate the robot, the server analyzes the information and provides appropriate operating instructions. For example, if a user asks, "How do I install a new part?", the server generates a response like this: "The installation steps are as follows: First, turn off the power and check the safety devices. Then install the part. Please let us know if you have any questions."
[0886] troubleshooting
[0887] When a user reports a problem with the robot, the server analyzes the cause and provides instructions based on emotion recognition to reduce stress and confusion. For example, if a user types, "My robot has stopped working, what should I do?", the server generates a response such as, "Please check the robot's power status and try restarting it. Is there anything else I can help you with?"
[0888] 4. Examples of prompts
[0889] "My robot has stopped working, what should I do?"
[0890] "How do I install the new parts?"
[0891] This allows the user to receive accurate work guidance in real time, improving the efficiency and safety of robot operation, and serves as a clear solution to the problem that the invention aims to solve.
[0892] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0893] Step 1:
[0894] The server hosts the chatbot application and loads the generative AI model and emotion recognition engine. The device installs and launches the chatbot application. The user interface is displayed, allowing the user to enter a message. When the user launches the application, the initial screen displays, "Hello. Is there anything I can help you with? How can I help you operate your robot?"
[0895] Step 2:
[0896] The user logs in to the application. The server checks the user information and retrieves the user profile. The retrieved user profile includes past operation history and setting information. The device displays a personalized UI based on the user's settings and past operation history. This allows the device to provide the user with appropriate suggestions and guidance.
[0897] Step 3:
[0898] The user enters questions or instructions into the chat box. For example, "How do I install a new part?" The system also analyzes the user's voice and facial expressions using an emotion engine. The device then sends these inputs to the server. The entered text data and voice data are then sent to the server.
[0899] Step 4:
[0900] The server uses a generative AI model to analyze natural language and an emotion recognition engine to recognize the user's emotional state. The analysis results identify the user's intentions, requests, and emotional state. For example, it may determine that the user wants to know the procedure for installing a new part. The user is feeling a little anxious. Natural language analysis extracts keywords from the input text data, and emotion recognition is performed by analyzing voice tone and facial expression data.
[0901] Step 5:
[0902] The server generates an appropriate response based on the analysis results and the user's emotional state. For example, it may describe the installation procedure in detail in a calm tone. It may generate a response such as, "The installation procedure is as follows: First, turn off the power and check the safety devices. Then, install the parts. Please let us know if you have any questions." The generated response is in text format.
[0903] Step 6:
[0904] The generated response is sent from the server to the terminal. The terminal displays the received response on the user interface. Based on the displayed message, the user can obtain specific instructions for performing the next action. For example, a message such as "The installation procedure is as follows. First, turn off the power and check the safety devices. Then install the parts. Please let us know if you have any questions" may be displayed.
[0905] Step 7:
[0906] When a user reports a specific problem, for example, "My robot has stopped working, what should I do?", the server analyzes the problem and evaluates the user's emotional state using an emotion recognition engine. The generated troubleshooting guidance is sent to the device as specific solution steps. The device displays the solution steps on the user interface and provides emotion-sensitive instructions. For example, a message such as "Please check the robot's power status and try restarting it. Is there anything else I can help you with?" is displayed.
[0907] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0908] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0909] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0910] [Third embodiment]
[0911] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0912] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0913] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0914] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0915] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0916] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0917] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0918] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0919] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0920] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0921] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0922] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0923] As an embodiment of the present invention, we provide a chatbot system that supports the use of operating systems (OS) for PCs and smartphones. The specific operation method and processing content of this system are described below.
[0924] First, the system includes means for receiving natural language input from a user, means for analyzing the received input, means for generating a response based on the analysis result, and means for displaying the response to the user.
[0925] Program processing overview
[0926] 1. Initialization and Setup
[0927] The server hosts the chatbot application, loads the generative AI model, and exposes API endpoints for interacting with the OS and users, starting the chatbot.
[0928] The device launches a dedicated chatbot app built into the OS, which displays a user interface (UI) and allows the user to enter a message.
[0929] The user accesses the chatbot application and starts a dialogue. The chatbot receives the natural language input from the user in text format.
[0930] 2. User Authentication and Profiling
[0931] A user logs in to an application. The server checks the user information and retrieves the user profile. The retrieved user profile includes past operation history and setting information. This allows personalized suggestions to be made to the user.
[0932] The device will prepare individually customized suggestions and guides based on the user's settings and past operation history.
[0933] 3. Natural Language Understanding and Dialogue Generation
[0934] Users enter questions or instructions into a chat box (e.g., "I want to install a new app"). The device sends this input to the server, which uses generative AI models to parse the natural language and generate an appropriate response.
[0935] The generated response is sent to the terminal and displayed on the user interface, allowing the user to intuitively understand the procedures for operating the OS.
[0936] 4. Function execution and operation support
[0937] Based on the response from the server, the device will take specific action. For example, if the user inputs "I want to install a new app," the server will search for the appropriate application and generate a response guiding the user through the installation process. The device will then display this response and guide them through the necessary operations.
[0938] The user follows the instructions to perform the action (e.g., clicks the install button). The device tracks the operation status in real time and displays additional guidance as needed.
[0939] 5. Troubleshooting
[0940] When a user reports a specific problem (e.g., "I can't connect to Wi-Fi"), the server analyzes the problem and identifies possible causes. The generated troubleshooting guidance is sent to the device as specific solution steps. The device displays the solution steps on the user interface and provides detailed instructions to the user.
[0941] Specific examples
[0942] New user onboarding
[0943] User: Starts up a newly purchased PC and uses the chatbot for the first time.
[0944] Device: Display "Nice to meet you. Let's go over the basic settings together. First, let's set the language."
[0945] User: Type "Please set to Japanese."
[0946] Terminal: Send this input to the server.
[0947] Server: Analyzes the message and generates instructions to change the language setting to Japanese.
[0948] Device: Change the language setting to Japanese and notify the user that "Settings are complete."
[0949] Introducing and explaining new features
[0950] User: After an update, launch the chatbot and type, "Tell me about the new features."
[0951] Server: Searches for information about new features and generates a summary of them.
[0952] On your device: "The new update includes a refreshed design, battery optimizations, and new security features."
[0953] Fault diagnosis and troubleshooting
[0954] User: Reports "No sound."
[0955] Server: Analyzes the problem, lists possible causes, and generates solutions.
[0956] Device: "Possible causes of no sound include the volume setting, muting, or speaker failure. First, check the volume setting." is displayed.
[0957] Users: Follow the instructions and double-check your volume settings.
[0958] In this way, a dedicated chatbot can assist users in various situations, making operating the OS more intuitive and efficient. This system allows users to effectively use the OS's functions and settings, and quickly respond to any problems that arise.
[0959] The processing flow will be explained below.
[0960] Step 1:
[0961] The server launches the chatbot application and loads the generative AI model, which is a machine learning model for natural language analysis and response generation. The server also opens an API endpoint and prepares to accept connection requests from devices.
[0962] Step 2:
[0963] The device installs and launches the chatbot application. The user interface (UI) is displayed, allowing the user to enter a message. The initial screen displays, "Nice to meet you. Welcome to the chatbot. How can I help you?"
[0964] Step 3:
[0965] The user enters login information (e.g., user ID and password) into the UI and clicks the "Login" button, which sends a login request from the terminal to the server.
[0966] Step 4:
[0967] The server compares the received login information with a database and performs authentication. If authentication is successful, the server obtains the user profile and sends it to the terminal. The user profile includes past operation history and individual setting information.
[0968] Step 5:
[0969] The device will then display a personalized UI based on the received user profile, such as shortcuts to frequently used functions and recently accessed applications.
[0970] Step 6:
[0971] The user types a question or instruction into the chat box (e.g., "I want to install a new app"), and the message is sent directly to the server.
[0972] Step 7:
[0973] The server receives the user's input message and uses a generative AI model to perform natural language analysis, identifying the user's intent and request (e.g., "I want to know the steps to install a new app").
[0974] Step 8:
[0975] The server generates an appropriate response based on the analysis results, such as a command to open the application store or guidance on installation procedures.
[0976] Step 9:
[0977] The generated response is sent from the server to the device, and may include specific instructions or options.
[0978] Step 10:
[0979] The device displays the received response on the user interface, for example, a confirmation message such as "Do you want to open the application store?"
[0980] Step 11:
[0981] The user follows the displayed instructions (e.g., answers "yes" and opens the application store), and the user's selection is sent back to the server.
[0982] Step 12:
[0983] Based on the user's selection, the server determines the next action and sends the necessary instructions to the terminal, such as to display search results for a particular application.
[0984] Step 13:
[0985] The device updates its user interface based on the instructions received from the server, for example by displaying a list of search results and allowing the user to select the app they want to install.
[0986] Step 14:
[0987] The user selects the app they want to install and sends an installation request to the server again.
[0988] Step 15:
[0989] The server confirms the installation request, initiates the necessary installation process, and notifies the device of the installation progress.
[0990] Step 16:
[0991] The terminal displays the installation progress to the user and notifies the user when the installation is complete.
[0992] Step 17:
[0993] The user launches the application after the installation is complete and begins using it.
[0994] Through this process, users can interact with the OS in natural language and smoothly perform the required actions.
[0995] Example 1
[0996] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0997] In recent years, the use of information devices such as PCs and smartphones has continued to expand, resulting in a demand for support regarding device operation and settings. However, conventional support systems have difficulty providing personalized support based on individual users' needs and operation history, limiting the improvement of usability. Furthermore, when it comes to troubleshooting, it is difficult to provide a prompt and accurate response to specific problems users face.
[0998] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0999] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, means for generating a corresponding response based on the analysis, means for displaying the response to the user, means for opening an API endpoint for communication with an OS and a user, means for loading a generative AI model, means for analyzing the user's input and generating an appropriate response, means for sending the generated response to a terminal, means for customizing a user interface, means for tracking user operations in real time and displaying additional guidance, and means for analyzing problems, identifying possible causes, and generating solutions. This enables personalized support based on the user's individual needs and operation history, and also enables quick and accurate troubleshooting.
[1000] "Natural language input from a user" refers to a message expressed in everyday language that is input by a user in the form of text, voice, or the like.
[1001] "Means for analyzing natural language input" refers to an algorithm or system that understands the natural language message received from the user and analyzes its grammatical structure and meaning.
[1002] "Means for generating a corresponding response" refers to a system or algorithm for generating an appropriate reply or instruction based on the results of natural language analysis.
[1003] "Means for displaying a response to a user" refers to an interface or device for displaying the generated response so that the user can visually confirm it.
[1004] "Means for opening API endpoints" refers to a system that provides the protocols and interfaces necessary to communicate with information terminals and operating systems.
[1005] "Means for loading a generative AI model" refers to the process or algorithm used to load a generative model into memory or a processor and make it executable.
[1006] "Means for analyzing user input and generating an appropriate response" refers to a process or system that analyzes user input and generates an optimal response based on the results.
[1007] The "means for transmitting the generated response to the terminal" refers to a system for sending the generated response data to the terminal via a network.
[1008] "Means for customizing a user interface" refers to a method or system for changing the design and functionality of an interface based on user settings and operation history.
[1009] "Means for tracking user operations in real time and displaying additional guidance" refers to a system that monitors the operations performed by a user in real time and provides next steps or additional instructions as needed.
[1010] "Means for analyzing problems, identifying possible causes, and generating solutions" refers to the systems and processes used to analyze user-reported problems, identify their causes, and generate appropriate solutions.
[1011] This invention relates to a chatbot system that supports the use of operating systems (OS) for PCs and smartphones. Specifically, it is a system that uses a generative AI model to analyze natural language input from users and generate appropriate responses.
[1012] This system mainly consists of three elements: a server, a terminal, and a user.
[1013] server
[1014] The server hosts the chatbot application and loads the generative AI model (e.g., GPT-4). The server exposes API endpoints for communication with the OS and users, and receives and analyzes user input. The server has the following capabilities:
[1015] A means of receiving natural language input from the user
[1016] A means of parsing natural language input
[1017] A means of generating a corresponding response based on the analysis results
[1018] A means of sending the generated response to the terminal
[1019] A means of analyzing problems, identifying possible causes, and generating solutions
[1020] Specifically, the server performs the following process.
[1021] 1. When receiving input from the user, retrieve the data through the appropriate API endpoint.
[1022] 2. The received input data is analyzed using a generative AI model, which performs tokenization and grammar analysis to understand meaning.
[1023] 3. Based on the analysis results, the server generates an appropriate response, referencing the user profile and past operation history to provide personalized suggestions.
[1024] 4. The generated response is sent to the device through the API endpoint.
[1025] Terminal
[1026] The device runs a dedicated chatbot application built into the operating system. The device is equipped with the following means:
[1027] A way to customize the user interface
[1028] A means to track user actions in real time and display additional guidance
[1029] Specifically, the terminal operates as follows.
[1030] 1. Launch the chatbot application and display a user interface that allows the user to enter a message.
[1031] 2. If there is any input from the user, send it to the server.
[1032] 3. Display the response received from the server in the user interface, providing interactive instructions as needed.
[1033] 4. Track user actions in real time and provide appropriate guidance when information is missing or additional guidance is needed.
[1034] User
[1035] Users access the chatbot application using a computer or smartphone and receive support through natural language dialogue. Users mainly perform the following operations:
[1036] 1. Log in to the application and provide the required authentication information.
[1037] 2. Enter questions or instructions in the chat box and receive responses from the server.
[1038] 3. Follow the displayed instructions and guides to operate the OS.
[1039] Specific examples
[1040] For example, when a user who has just purchased a new PC performs initial setup, the following dialogue may take place:
[1041] User: "I just started up my new PC. What do I do now?"
[1042] Server: "Nice to meet you. Let's go over some basic setup together. First, let's set up the language."
[1043] User: "Please set it to Japanese."
[1044] Server: "The language setting has been changed to Japanese. Next, set the time zone."
[1045] Also, if the user wants to install a new application, the prompt sentence is as follows:
[1046] Prompt for "I want to install a new app":
[1047] Input to the server: "The user wants to install a new app. Please guide them through the installation process."
[1048] Example response: "Find the right application and guide me through the installation process."
[1049] This invention allows users to intuitively and efficiently use OS functions and settings, and also allows them to receive prompt and accurate support when problems occur.
[1050] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1051] Step 1: Initialization and Setup
[1052] server:
[1053] The server hosts the chatbot application, loads the generative AI model (e.g., GPT-4), opens an API endpoint, and prepares to communicate with the OS and the user. Specific operations include the following steps:
[1054] The server reads specific configuration files and database connection information.
[1055] The server loads the generative AI model into memory and makes it executable.
[1056] Device:
[1057] The device launches a dedicated chatbot app and displays a user interface (UI), allowing the user to enter a message. Specifically, the process involves the following steps:
[1058] The device prepares the UI for display, rendering text boxes, submit buttons, etc.
[1059] The terminal checks the connection to the server and whether communication is possible.
[1060] User:
[1061] A user accesses a chatbot application and starts a conversation. The specific operations include the following steps:
[1062] The user checks the application's UI and enters an initial message such as "Hello."
[1063] Step 2: User authentication and profiling
[1064] User:
[1065] A user logs into an application, typically by entering a username and password, and sometimes by using biometric authentication such as facial recognition or fingerprint authentication.
[1066] server:
[1067] The server checks the user's authentication information against the database and retrieves the user profile. Specifically, the following process is performed:
[1068] The server takes the entered username and password and issues an SQL query to the database to verify it.
[1069] If the authentication is successful, the server obtains the user profile (including operation history and setting information) and stores it in memory.
[1070] Device:
[1071] The device customizes the UI based on the user profile information received from the server. Specifically, the following process is performed:
[1072] The device changes the UI style based on the user's language and theme settings.
[1073] Step 3: Natural Language Understanding and Dialogue Generation
[1074] User:
[1075] The user types a question or instruction into the chat box, for example, "I want to install a new app."
[1076] Device:
[1077] The terminal sends the user's input to the server. The specific procedure is as follows:
[1078] The device receives the input as text data and sends an HTTP POST request through the API endpoint.
[1079] server:
[1080] The server analyzes the received user input using a generative AI model. Specifically, the following data processing and calculations are performed:
[1081] The server tokenizes the input data, performs grammatical analysis, and analyzes it to understand its meaning.
[1082] Based on the analysis results, an appropriate response is generated.
[1083] Device:
[1084] The generated response is sent to the terminal and displayed on the user interface. Specifically, the process includes the following steps:
[1085] The terminal acquires the response data received from the server and displays it in the text box.
[1086] Step 4: Implementing functions and assisting operations
[1087] server:
[1088] The server then generates a response that recommends a specific action, such as instructions for installing a new app.
[1089] Device:
[1090] The terminal guides the user according to the procedure from the server. Specifically, the terminal performs the following procedure.
[1091] The device will display instructions such as "Click the install button."
[1092] User:
[1093] The user follows the instructions displayed on the device, for example, "click the install button."
[1094] Step 5: Troubleshooting
[1095] User:
[1096] The user reports a specific problem, for example, "I can't connect to Wi-Fi."
[1097] server:
[1098] The server analyzes the problem, identifies possible causes, and generates solutions. Specifically, the following data processing and calculations are performed:
[1099] The server references the data for troubleshooting and identifies the cause through natural language analysis.
[1100] The server generates solutions and organizes them into guidance.
[1101] Device:
[1102] The generated troubleshooting guidance is displayed in the user interface, specifically by following the steps below.
[1103] The device will display instructions such as "First, please check your volume settings."
[1104] This system allows users to operate the OS and solve problems intuitively and efficiently.
[1105] (Application example 1)
[1106] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1107] Conventional systems make it difficult for users to operate and configure the system in natural language, and they lack sufficient support for security-related settings and troubleshooting. Furthermore, they do not provide personalized suggestions that effectively utilize user profiles and operation histories, making it difficult for users to operate the system intuitively and efficiently.
[1108] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1109] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, and means for generating a corresponding response based on the analysis, thereby enabling the user to perform security settings and other operations in natural language.
[1110] Furthermore, the server includes means for performing specific actions based on the generated responses and means for analyzing problems and generating solutions during troubleshooting, thereby enabling the user to obtain specific troubleshooting and operation guidance.
[1111] Furthermore, by including a means for acquiring and managing a user profile, a means for providing personalized suggestions to the user based on the profile, and a means for adjusting security settings based on the user's operation history and setting information, intuitive and individually optimized operation becomes possible, allowing the user to efficiently manage security settings.
[1112] "Natural language input" is an input format in which a user gives instructions or asks questions to a system in everyday language.
[1113] "Analysis" refers to the process of converting input natural language into a form that a machine can understand and interpreting its meaning.
[1114] A "response" is a reply or instruction from the system that is generated based on the parsed natural language input.
[1115] A "user profile" is a data set that collects information about a user, including individual settings and past operation history.
[1116] "Personalized suggestions" are suggestions that individually provide users with information and guidance that is most suitable for them based on their user profile.
[1117] The "operation history" is a record of a series of operations and setting changes that the user has performed in the past.
[1118] A "specific action" is a specific operation or process that the system actually performs based on a user's instruction.
[1119] "Troubleshooting" is the process of identifying the cause of a problem and providing a solution.
[1120] A "generative AI model" is an artificial intelligence technology that uses machine learning to understand natural language and generate appropriate responses.
[1121] "Security settings" refers to a set of settings and policies that ensure the security of a device or system.
[1122] As an embodiment of the present invention, we will explain a chatbot system in which a server and a terminal cooperate to assist users in security configuration and troubleshooting. This system utilizes a generative AI model to provide appropriate responses to security-related instructions and questions entered by the user in natural language.
[1123] Programs and hardware / software used
[1124] An API (specifically, the OpenAI API) that runs the generative AI model is installed on the server, which analyzes input from users and generates responses. A smartphone is used as the terminal, and a chatbot app is installed on this smartphone. Users interact with the system through this app.
[1125] Natural language input processing begins when a user enters instructions or questions in natural language into the chatbot app. This input is sent to the server as text data, where it is analyzed by the generative AI model. Based on the results of the analysis, the server generates an appropriate response and sends it back to the device.
[1126] User profile management is achieved by applications sending user information to a server, which then stores the user's operation history and settings information and generates personalized suggestions and actions based on this information, which can be useful when adjusting security settings.
[1127] For example, if a user types "check Wi-Fi settings," the device will execute a script based on the generated response to help check the actual Wi-Fi settings, allowing users to easily troubleshoot security settings and resolve issues.
[1128] When troubleshooting, users report their issues in natural language and their input is sent to the server, which uses generative AI models to analyze the cause of the problem and generate a solution as a response. For example, if a user reports "no sound," specific steps such as "check your volume settings" are provided.
[1129] Specific examples
[1130] User: "Check the security status of your device"
[1131] Chatbot: "We're running a security check on your device. Please wait a moment... All security settings are OK."
[1132] Example prompt sentence:
[1133] User input: Check the security status of your device
[1134] Security Chatbot Response: We're going to perform a security check on your device. Please wait... All your security settings are fine.
[1135] In this way, users can easily perform security-related operations in natural language through the chatbot and receive appropriate troubleshooting guidance when problems arise, enabling users to manage their security settings intuitively and efficiently.
[1136] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1137] Step 1:
[1138] The device launches the chatbot app, and the user inputs instructions or questions in natural language. The input text data is sent to the server via an API call. The server receives the text data and prepares it for analysis by the generative AI model.
[1139] Input: User's natural language input (text data)
[1140] Output: Text data sent to the server
[1141] Step 2:
[1142] The server passes the received natural language input to the generative AI model, which begins the analysis process. The generative AI model analyzes the input natural language, understands its meaning, and generates an appropriate response. The analysis uses the OpenAI API.
[1143] Input: Text data sent to the server
[1144] Output: Analysis results from the generative AI model
[1145] Step 3:
[1146] Based on the generated analysis, the server generates a specific response to the user's instruction, such as "check your Wi-Fi settings" or any other troubleshooting steps required.
[1147] Input: Analysis results from generative AI model
[1148] Output: Response data to the user
[1149] Step 4:
[1150] The server sends the generated response data to the terminal, which receives the response and displays it on the user interface within the chatbot app. The user can then view it and take the next action.
[1151] Input: Response data to the user
[1152] Output: Response data sent to the terminal
[1153] Step 5:
[1154] The user takes action based on the generated response. For example, if the instruction is "Check your Wi-Fi settings," the user opens the Wi-Fi settings screen and adjusts the settings. After completing this action, the user again provides feedback to the chatbot.
[1155] Input: Response data displayed to the user
[1156] Output: User's actual actions
[1157] Step 6:
[1158] If the user wants to input additional instructions to resolve the issue, the process repeats from step 1. For example, if the user inputs "Please also check the volume settings," the server will analyze it again and return the appropriate instructions.
[1159] Input: Additional user instructions (natural language input)
[1160] Output: The analysis process is restarted.
[1161] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1162] As an embodiment of the present invention, we provide a chatbot system that supports the use of operating systems (OS) for PCs and smartphones. In particular, we will explain an embodiment that incorporates an engine that recognizes the user's emotions.
[1163] First, the system has the means to receive natural language input from users, parse it, generate and display a response, and, combined with an emotion engine, recognizes the user's emotional state and adapts responses based on that, achieving a high level of personalization.
[1164] Program processing overview
[1165] 1. Initialization and Setup
[1166] The server hosts the chatbot application and loads the generative AI model, a machine learning model for natural language analysis and response generation, and the emotion engine, a model for recognizing user emotions.
[1167] The chatbot application is installed on the device and launched. The user interface (UI) is displayed, allowing the user to enter a message. The initial screen displays, "Hello. Is there anything I can help you with?"
[1168] The user accesses the chatbot application and starts a dialogue, receiving the user's natural language or voice input in text format.
[1169] 2. User Authentication and Profiling
[1170] A user logs in to an application. The server checks the user information and retrieves the user profile. The retrieved user profile includes past operation history and setting information.
[1171] The device displays a personalized UI based on the user's settings and past operation history, allowing it to provide more appropriate suggestions and guidance to the user.
[1172] 3. Natural Language Understanding and Emotion Recognition
[1173] Users enter questions or instructions into the chat box (e.g., "I want to install a new app"). The system also analyzes the user's facial expressions and tone of voice using an emotion engine. The device then sends these inputs to the server.
[1174] The server uses a generative AI model to analyze natural language and an emotion engine to recognize the user's emotions. The analysis results identify the user's intent, request, and emotional state (e.g., "I want to know the steps to install a new app. The user is feeling a bit anxious and worried.").
[1175] 4. Response generation and emotional adaptation
[1176] Based on the analysis and emotional state, the server generates an appropriate response, such as a command to open the application store or friendly, gentle installation instructions.
[1177] The generated response is sent from the server to the device, and may include specific instructions or an emotionally sensitive tone.
[1178] The device displays the received response on the user interface, for example, a message such as "Open the application store and install the desired app. If you have any questions, please feel free to ask."
[1179] 5. Troubleshooting
[1180] When a user reports a specific problem (e.g., "I can't connect to Wi-Fi"), the server analyzes the problem and evaluates the user's emotional state using an emotion engine. The generated troubleshooting guidance is sent to the device as specific solution steps. The device then displays the solution steps on the user interface and provides emotion-sensitive instructions.
[1181] Specific examples
[1182] New user onboarding
[1183] User: Starts up a newly purchased PC and uses the chatbot for the first time.
[1184] Device: Display "Hello. Let's go through the basic setup together. If you have any questions, please feel free to ask."
[1185] User: Type "Please set to Japanese."
[1186] Terminal: Send this input to the server.
[1187] Server: Parses the message and generates instructions to change the language setting to Japanese. Recognizes that the user is feeling a bit anxious and generates a response in a friendly tone.
[1188] Device: Change the language setting to Japanese and notify the user, "Setup complete. You can now use Japanese. Is there anything else we can help you with?"
[1189] Introducing and explaining new features
[1190] User: After an update, launch the chatbot and type, "Tell me about the new features."
[1191] Server: Searches for information about new features and generates summaries of them, recognizing the user's emotional state and crafting interesting descriptions.
[1192] On your device: "Our new update features a fresh design, battery optimizations, and new security features."
[1193] Fault diagnosis and troubleshooting
[1194] User: Reports "No sound."
[1195] Server: Analyzes the problem and recognizes the user's emotional state using an emotion engine. Lists possible causes and generates solutions.
[1196] Device: "Possible reasons for the lack of sound include the volume setting, muting, or a faulty speaker. First, check the volume setting. If you have any questions, please let us know."
[1197] User: Follow the instructions and check the volume settings again. In this way, a system incorporating an emotion engine allows users to receive responses that take their emotions into consideration, allowing them to operate the OS with greater peace of mind.
[1198] The processing flow will be explained below.
[1199] Step 1:
[1200] The server launches the chatbot application and loads the generative AI model and emotion engine. The generative AI model is a machine learning model for natural language analysis and response generation, and the emotion engine is a model for recognizing user emotions. The server also opens an API endpoint and prepares to accept connection requests from devices.
[1201] Step 2:
[1202] The chatbot application is installed on the device and launched. The user interface (UI) is displayed, allowing the user to enter a message. The initial screen displays, "Hello. Is there anything I can help you with?"
[1203] Step 3:
[1204] The user enters login information (e.g., user ID and password) into the UI and clicks the "Login" button, which sends a login request from the terminal to the server.
[1205] Step 4:
[1206] The server compares the received login information with a database and performs authentication. If authentication is successful, the server obtains the user profile and sends it to the terminal. The user profile includes past operation history and individual setting information.
[1207] Step 5:
[1208] The device will then display a personalized UI based on the received user profile, such as shortcuts to frequently used functions and recently accessed applications.
[1209] Step 6:
[1210] Users enter questions or instructions into the chat box (e.g., "I want to install a new app"). The entered message is sent directly to the server. The system also analyzes the user's facial expressions and tone of voice using an emotion engine.
[1211] Step 7:
[1212] The server receives the user's input message and uses a generative AI model to perform natural language analysis, identifying the user's intent and request (e.g., "I want to know the steps to install a new app").
[1213] Step 8:
[1214] The server uses an emotion engine to recognize the user's emotional state, for example, to determine whether the user is feeling anxious or worried.
[1215] Step 9:
[1216] The server generates an appropriate response based on the analysis and emotional state, for example, "Do you want to open the app store?" in a calm and reassuring tone.
[1217] Step 10:
[1218] The generated response is sent from the server to the device, and may include specific instructions or an emotionally sensitive tone.
[1219] Step 11:
[1220] The device displays the received response on the user interface, for example, "Open the application store and install the desired app. If you have any questions, please feel free to ask."
[1221] Step 12:
[1222] The user follows the displayed instructions (e.g., answers "yes" and opens the application store), and the user's selection is sent back to the server.
[1223] Step 13:
[1224] Based on the user's selection, the server determines the next action and sends the necessary instructions to the terminal, such as to display search results for a particular application.
[1225] Step 14:
[1226] The device updates its user interface based on the instructions received from the server, for example by displaying a list of search results and allowing the user to select the app they want to install.
[1227] Step 15:
[1228] The user selects the app they want to install and sends an installation request to the server again.
[1229] Step 16:
[1230] The server confirms the installation request, initiates the necessary installation process, and notifies the device of the installation progress.
[1231] Step 17:
[1232] The terminal displays the installation progress to the user and notifies the user when the installation is complete.
[1233] Step 18:
[1234] Once the installation is complete, the user launches the application and begins using it. Through this process, the user can interact with the OS based on natural language and emotions to smoothly perform the required actions.
[1235] Example 2
[1236] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1237] Conventional chatbot systems have difficulty generating appropriate responses to users' natural language input, and in particular, they are unable to recognize emotions and reflect them in responses, which means they are unable to provide a satisfying conversational experience for users.In addition, they lack personalized suggestions that utilize the user's profile information and past operation history, making it difficult to achieve user-friendly operation.
[1238] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1239] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, means for generating a corresponding response based on the analysis, means for displaying the response to the user, means for analyzing natural language using a generative AI model, means for identifying the user's emotional state using an emotion recognition engine, means for generating an appropriate response based on the analysis result and the emotional state, and means for displaying the response in a tone that takes the user's emotions into consideration. This makes it possible to provide more personalized suggestions and guidance while taking the user's emotions into consideration, thereby achieving a highly satisfying dialogue experience.
[1240] "Natural language input" refers to text or voice input entered by a user in the form of everyday spoken language.
[1241] A "generative AI model" refers to an algorithmic model that uses machine learning technology to analyze natural language and generate responses.
[1242] "Emotion recognition engine" refers to a software module for analyzing and identifying an emotional state from a user's input.
[1243] A "user profile" refers to a collection of data including a user's personal information, past operation history, and setting information.
[1244] "Personalization" refers to providing customized suggestions and responses based on individual user characteristics and past behavior.
[1245] "Operation history" refers to a record of operations that a user has performed in the past.
[1246] "Setting information" refers to information about various settings that a user makes when using an application or system.
[1247] A "server" refers to a computer or software system used to receive requests from users and process or respond to data.
[1248] "Terminal" refers to the device that the user actually operates (e.g., a PC or smartphone).
[1249] "User interface" refers to the screen display and input means that allow a user to interact with a system.
[1250] "Login" refers to the authentication procedure required for a user to access a system.
[1251] "Analysis" refers to the processing of data to understand a user's natural language input and emotional state.
[1252] A "response" refers to a reply or instruction that a system generates and provides in response to a user's input.
[1253] As an embodiment of the present invention, we provide a chatbot system that supports the use of operating systems (OS) for personal computers and smartphones. This system incorporates a generative AI model and an emotion recognition engine.
[1254] First, the server hosts the chatbot application and loads a generative AI model (e.g., GPT-3.5) and an emotion recognition engine (e.g., a popular emotion recognition API). The generative AI model is used to parse natural language and generate responses, and the emotion recognition engine is used to recognize the user's emotional state. These models and engines are initialized by the server by loading the API key and model data.
[1255] The device then allows the user to install and launch the chatbot application. When the user launches the application, a user interface appears, displaying a welcome message that reads, "Hello, how can I help you?"
[1256] The user accesses the chatbot application and starts a dialogue using natural language or voice. The input content is received in text format and sent from the terminal to the server. If the user inputs voice, the voice is converted into text and sent to the server.
[1257] When a user logs in to an application, the server checks the user information and obtains profile data (past operation history, setting information, etc.). Based on the obtained profile data, a personalized UI is displayed to the user. For example, frequently used functions can be displayed as shortcuts based on the past operation history.
[1258] The server analyzes the natural language input from the user using a generative AI model to identify the user's intentions and requests. It also uses an emotion recognition engine to recognize the user's emotional state. For example, if a user inputs "I want to install a new app," the analysis results might determine that "I want to know the steps to install a new app, and I'm feeling a bit impatient and anxious."
[1259] Based on the analysis results and the emotional state, the server uses a generative AI model to generate an appropriate response, such as "Please visit the application store and install the app you want. If you have any questions, please feel free to ask." This response is generated in an emotionally sensitive tone.
[1260] The generated response is sent to the device and displayed in the user interface. The displayed response is provided in an emotionally sensitive tone, allowing the user to proceed with the operation with confidence.
[1261] When a user reports a specific problem, for example, "I can't connect to Wi-Fi," the server analyzes the problem and uses an emotion recognition engine to assess the user's emotional state. Based on the analysis, specific solutions are generated and sent to the device. For example, a message might read, "Possible causes of the lack of audio include volume settings, muting, or a faulty speaker. First, check your volume settings."
[1262] Specific examples
[1263] New user onboarding
[1264] User: Starts up a newly purchased PC and uses the chatbot for the first time.
[1265] Device: Display "Hello. Let's go through the basic setup together. If you have any questions, please feel free to ask."
[1266] User: Type "Please set to Japanese."
[1267] Terminal: Send this input to the server.
[1268] Server: Parses the message and generates instructions to change the language setting to Japanese. Recognizes that the user is feeling a bit anxious and generates a response in a friendly tone.
[1269] Device: Change the language setting to Japanese and notify the user, "Setup complete. You can now use Japanese. Is there anything else we can help you with?"
[1270] Introducing and explaining new features
[1271] User: After an update, launch the chatbot and type, "Tell me about the new features."
[1272] Server: Searches for information about new features and generates summaries of them, recognizing the user's emotional state and crafting interesting descriptions.
[1273] On your device: "Our new update features a fresh design, battery optimizations, and new security features."
[1274] Fault diagnosis and troubleshooting
[1275] User: Reports "No sound."
[1276] Server: Analyzes the problem and recognizes the user's emotional state using an emotion recognition engine. Lists possible causes and generates solutions.
[1277] Device: "Possible reasons for the lack of sound include the volume setting, muting, or a faulty speaker. First, check the volume setting. If you have any questions, please let us know."
[1278] In this way, a system incorporating an emotion recognition engine allows users to receive responses that take their emotions into consideration, allowing them to operate the OS with greater peace of mind.
[1279] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1280] Step 1:
[1281] The server hosts the chatbot application and loads the generative AI model (e.g., GPT-3.5) and emotion recognition engine (e.g., a popular emotion recognition API). Specifically, it loads the API key and model data to initialize the model. The input at this stage is the required API key and model data, and the output is a ready-to-use generative AI model and emotion recognition engine.
[1282] Step 2:
[1283] When a user installs and launches a chatbot application, the device loads the user interface. The initial screen displays "Hello, is there anything I can help you with?" When a user launches the application, the input is the application launch command, and the output is a user interface displaying the message.
[1284] Step 3:
[1285] The user inputs questions or instructions to the chatbot using natural language or voice. For example, they might input "I want to install a new app." This input is received by the device in text format. If the input is voice, the voice is converted to text. At this stage, the input is the user's natural language or voice input, and the output is text data.
[1286] Step 4:
[1287] The device sends the received text data to the server. It also sends text data converted from speech. Specifically, it sends the text data to the server via an HTTP request. At this stage, the input is user data in text format, and the output is the data sent to the server.
[1288] Step 5:
[1289] The server analyzes the received text data using a generative AI model to identify the user's intentions and requests. It also uses an emotion recognition engine to recognize the user's emotional state. For example, if a user inputs "I want to install a new app," the analysis results might determine that "I want to know the steps to install a new app. The user is feeling a bit impatient and anxious." The input at this stage is text data, and the output is the analysis results regarding the user's intentions and emotional state.
[1290] Step 6:
[1291] The server uses a generative AI model to generate an appropriate response based on the analysis results and the emotional state. For example, it generates a message such as, "Open the application store and install the app you want. If you have any questions, please feel free to ask." This response is generated in an emotionally sensitive tone. The input at this stage is the analysis results and the emotional state data, and the output is an appropriate response message.
[1292] Step 7:
[1293] The server sends the generated response message to the device. The device receives the response and displays it on its user interface. The displayed message contains specific instructions such as "Open the application store and install the desired app." The input at this stage is the response message, and the output is the message displayed on the user interface.
[1294] Step 8:
[1295] When a user reports a specific problem, for example, "I can't connect to Wi-Fi," the server analyzes the problem and uses an emotion recognition engine to assess the user's emotional state. The analysis results in a list of possible causes and generates a solution. For example, a message might read, "Possible causes of the no sound include the volume setting, muting, or a faulty speaker. Please check the volume setting first." The input at this stage is text data about the specific problem, and the output is the problem analysis and a solution message.
[1296] Step 9:
[1297] The terminal displays the solution message received from the server on the user interface. For example, it displays "Please check your volume settings." This message is expressed in an emotionally sensitive tone. The input at this stage is the solution message, and the output is the solution message displayed on the user interface.
[1298] This allows users to receive emotionally sensitive responses and operate with peace of mind.The system improves the user experience by combining natural language input analysis and emotion recognition.
[1299] (Application example 2)
[1300] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1301] In factory work environments, workers are required to be able to operate robots and troubleshoot them quickly and reliably. However, conventional systems have difficulty responding to the user's emotional state and real-time work environment, which often causes stress and confusion for workers. This can hinder efficient work execution and reduce safety.
[1302] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1303] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, means for generating a corresponding response, means for recognizing the user's emotional state, and means for adapting the response based on the user's emotional state. This enables generation of a response that takes the user's emotions into consideration. The server also includes means for acquiring the user's operation history and setting information, means for providing information related to the work environment in real time, and means for providing appropriate operation guidance to the user. This enables appropriate guidance based on real-time information about the work environment, allowing work to be performed safely and efficiently.
[1304] "Natural language input" refers to a language that is input by a user in the form of text, speech, or other language that is commonly used by humans.
[1305] "Analysis" refers to the syntactic and semantic processing and interpretation of input natural language.
[1306] A "response" is a response to a question or instruction that is generated based on the results of the analysis.
[1307] "Display" refers to providing the generated response to the user visually or audibly.
[1308] "Emotional state" refers to the user's psychological and emotional state as determined by their tone of voice and facial expression.
[1309] A "profile" is a collection of data about a user, including the user's personal information, operation history, setting information, and the like.
[1310] "Personalization" refers to providing individualized offers and information based on a user's profile information.
[1311] An "operation history" is a record of operations and actions that a user has performed in the past.
[1312] "Setting information" refers to setting items and parameters that are individually set by the user for the system or application.
[1313] "Work environment information" is information relating to the user's real-time work location and situation.
[1314] "Operational guidance" refers to procedures or instructions provided to a user to perform a specific task or operation.
[1315] This embodiment of the present invention is a chatbot system with advanced response capabilities incorporating user emotion recognition. It is particularly intended for supporting robotic operation in factory work environments. The system analyzes the user's natural language input, generates corresponding responses, and displays those responses. It also provides appropriate operational guidance in real time based on the user profile and work environment information.
[1316] 1. Hardware and Software Used
[1317] This system uses the following hardware and software:
[1318] Hardware: Smartphones, smart glasses, head-mounted displays
[1319] software:
[1320] Python: To execute the main logic of the program
[1321] OpenAI API: Natural Language Analysis and Response Generation
[1322] Emotion Recognition API: To recognize user emotions
[1323] 2. Program processing content
[1324] The server uses a generative AI model to analyze natural language and generate appropriate responses to user input. Furthermore, it uses the Emotion Recognition API to determine the user's emotional state and adapts responses based on that emotion. Specific responses include:
[1325] 3. System Operation Example
[1326] Operation Guidance
[1327] When a user inquires about how to operate the robot, the server analyzes the information and provides appropriate operating instructions. For example, if a user asks, "How do I install a new part?", the server generates a response like this: "The installation steps are as follows: First, turn off the power and check the safety devices. Then install the part. Please let us know if you have any questions."
[1328] troubleshooting
[1329] When a user reports a problem with the robot, the server analyzes the cause and provides instructions based on emotion recognition to reduce stress and confusion. For example, if a user types, "My robot has stopped working, what should I do?", the server generates a response such as, "Please check the robot's power status and try restarting it. Is there anything else I can help you with?"
[1330] 4. Examples of prompts
[1331] "My robot has stopped working, what should I do?"
[1332] "How do I install the new parts?"
[1333] This allows the user to receive accurate work guidance in real time, improving the efficiency and safety of robot operation, and serves as a clear solution to the problem that the invention aims to solve.
[1334] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1335] Step 1:
[1336] The server hosts the chatbot application and loads the generative AI model and emotion recognition engine. The device installs and launches the chatbot application. The user interface is displayed, allowing the user to enter a message. When the user launches the application, the initial screen displays, "Hello. Is there anything I can help you with? How can I help you operate your robot?"
[1337] Step 2:
[1338] The user logs in to the application. The server checks the user information and retrieves the user profile. The retrieved user profile includes past operation history and setting information. The device displays a personalized UI based on the user's settings and past operation history. This allows the device to provide the user with appropriate suggestions and guidance.
[1339] Step 3:
[1340] The user enters questions or instructions into the chat box. For example, "How do I install a new part?" The system also analyzes the user's voice and facial expressions using an emotion engine. The device then sends these inputs to the server. The entered text data and voice data are then sent to the server.
[1341] Step 4:
[1342] The server uses a generative AI model to analyze natural language and an emotion recognition engine to recognize the user's emotional state. The analysis results identify the user's intentions, requests, and emotional state. For example, it may determine that the user wants to know the procedure for installing a new part. The user is feeling a little anxious. Natural language analysis extracts keywords from the input text data, and emotion recognition is performed by analyzing voice tone and facial expression data.
[1343] Step 5:
[1344] The server generates an appropriate response based on the analysis results and the user's emotional state. For example, it may describe the installation procedure in detail in a calm tone. It may generate a response such as, "The installation procedure is as follows: First, turn off the power and check the safety devices. Then, install the parts. Please let us know if you have any questions." The generated response is in text format.
[1345] Step 6:
[1346] The generated response is sent from the server to the terminal. The terminal displays the received response on the user interface. Based on the displayed message, the user can obtain specific instructions for performing the next action. For example, a message such as "The installation procedure is as follows. First, turn off the power and check the safety devices. Then install the parts. Please let us know if you have any questions" may be displayed.
[1347] Step 7:
[1348] When a user reports a specific problem, for example, "My robot has stopped working, what should I do?", the server analyzes the problem and evaluates the user's emotional state using an emotion recognition engine. The generated troubleshooting guidance is sent to the device as specific solution steps. The device displays the solution steps on the user interface and provides emotion-sensitive instructions. For example, a message such as "Please check the robot's power status and try restarting it. Is there anything else I can help you with?" is displayed.
[1349] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1350] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1351] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1352] [Fourth embodiment]
[1353] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1354] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1355] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1356] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1357] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1358] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1359] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1360] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1361] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1362] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1363] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1364] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1365] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1366] As an embodiment of the present invention, we provide a chatbot system that supports the use of operating systems (OS) for PCs and smartphones. The specific operation method and processing content of this system are described below.
[1367] First, the system includes means for receiving natural language input from a user, means for analyzing the received input, means for generating a response based on the analysis result, and means for displaying the response to the user.
[1368] Program processing overview
[1369] 1. Initialization and Setup
[1370] The server hosts the chatbot application, loads the generative AI model, and exposes API endpoints for interacting with the OS and users, starting the chatbot.
[1371] The device launches a dedicated chatbot app built into the OS, which displays a user interface (UI) and allows the user to enter a message.
[1372] The user accesses the chatbot application and starts a dialogue. The chatbot receives the natural language input from the user in text format.
[1373] 2. User Authentication and Profiling
[1374] A user logs in to an application. The server checks the user information and retrieves the user profile. The retrieved user profile includes past operation history and setting information. This allows personalized suggestions to be made to the user.
[1375] The device will prepare individually customized suggestions and guides based on the user's settings and past operation history.
[1376] 3. Natural Language Understanding and Dialogue Generation
[1377] Users enter questions or instructions into a chat box (e.g., "I want to install a new app"). The device sends this input to the server, which uses generative AI models to parse the natural language and generate an appropriate response.
[1378] The generated response is sent to the terminal and displayed on the user interface, allowing the user to intuitively understand the procedures for operating the OS.
[1379] 4. Function execution and operation support
[1380] Based on the response from the server, the device will take specific action. For example, if the user inputs "I want to install a new app," the server will search for the appropriate application and generate a response guiding the user through the installation process. The device will then display this response and guide them through the necessary operations.
[1381] The user follows the instructions to perform the action (e.g., clicks the install button). The device tracks the operation status in real time and displays additional guidance as needed.
[1382] 5. Troubleshooting
[1383] When a user reports a specific problem (e.g., "I can't connect to Wi-Fi"), the server analyzes the problem and identifies possible causes. The generated troubleshooting guidance is sent to the device as specific solution steps. The device displays the solution steps on the user interface and provides detailed instructions to the user.
[1384] Specific examples
[1385] New user onboarding
[1386] User: Starts up a newly purchased PC and uses the chatbot for the first time.
[1387] Device: Display "Nice to meet you. Let's go over the basic settings together. First, let's set the language."
[1388] User: Type "Please set to Japanese."
[1389] Terminal: Send this input to the server.
[1390] Server: Analyzes the message and generates instructions to change the language setting to Japanese.
[1391] Device: Change the language setting to Japanese and notify the user that "Settings are complete."
[1392] Introducing and explaining new features
[1393] User: After an update, launch the chatbot and type, "Tell me about the new features."
[1394] Server: Searches for information about new features and generates a summary of them.
[1395] On your device: "The new update includes a refreshed design, battery optimizations, and new security features."
[1396] Fault diagnosis and troubleshooting
[1397] User: Reports "No sound."
[1398] Server: Analyzes the problem, lists possible causes, and generates solutions.
[1399] Device: "Possible causes of no sound include the volume setting, muting, or speaker failure. First, check the volume setting." is displayed.
[1400] Users: Follow the instructions and double-check your volume settings.
[1401] In this way, a dedicated chatbot can assist users in various situations, making operating the OS more intuitive and efficient. This system allows users to effectively use the OS's functions and settings, and quickly respond to any problems that arise.
[1402] The processing flow will be explained below.
[1403] Step 1:
[1404] The server launches the chatbot application and loads the generative AI model, which is a machine learning model for natural language analysis and response generation. The server also opens an API endpoint and prepares to accept connection requests from devices.
[1405] Step 2:
[1406] The device installs and launches the chatbot application. The user interface (UI) is displayed, allowing the user to enter a message. The initial screen displays, "Nice to meet you. Welcome to the chatbot. How can I help you?"
[1407] Step 3:
[1408] The user enters login information (e.g., user ID and password) into the UI and clicks the "Login" button, which sends a login request from the terminal to the server.
[1409] Step 4:
[1410] The server compares the received login information with a database and performs authentication. If authentication is successful, the server obtains the user profile and sends it to the terminal. The user profile includes past operation history and individual setting information.
[1411] Step 5:
[1412] The device will then display a personalized UI based on the received user profile, such as shortcuts to frequently used functions and recently accessed applications.
[1413] Step 6:
[1414] The user types a question or instruction into the chat box (e.g., "I want to install a new app"), and the message is sent directly to the server.
[1415] Step 7:
[1416] The server receives the user's input message and uses a generative AI model to perform natural language analysis, identifying the user's intent and request (e.g., "I want to know the steps to install a new app").
[1417] Step 8:
[1418] The server generates an appropriate response based on the analysis results, such as a command to open the application store or guidance on installation procedures.
[1419] Step 9:
[1420] The generated response is sent from the server to the device, and may include specific instructions or options.
[1421] Step 10:
[1422] The device displays the received response on the user interface, for example, a confirmation message such as "Do you want to open the application store?"
[1423] Step 11:
[1424] The user follows the displayed instructions (e.g., answers "yes" and opens the application store), and the user's selection is sent back to the server.
[1425] Step 12:
[1426] Based on the user's selection, the server determines the next action and sends the necessary instructions to the terminal, such as to display search results for a particular application.
[1427] Step 13:
[1428] The device updates its user interface based on the instructions received from the server, for example by displaying a list of search results and allowing the user to select the app they want to install.
[1429] Step 14:
[1430] The user selects the app they want to install and sends an installation request to the server again.
[1431] Step 15:
[1432] The server confirms the installation request, initiates the necessary installation process, and notifies the device of the installation progress.
[1433] Step 16:
[1434] The terminal displays the installation progress to the user and notifies the user when the installation is complete.
[1435] Step 17:
[1436] The user launches the application after the installation is complete and begins using it.
[1437] Through this process, users can interact with the OS in natural language and smoothly perform the required actions.
[1438] Example 1
[1439] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1440] In recent years, the use of information devices such as PCs and smartphones has continued to expand, resulting in a demand for support regarding device operation and settings. However, conventional support systems have difficulty providing personalized support based on individual users' needs and operation history, limiting the improvement of usability. Furthermore, when it comes to troubleshooting, it is difficult to provide a prompt and accurate response to specific problems users face.
[1441] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1442] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, means for generating a corresponding response based on the analysis, means for displaying the response to the user, means for opening an API endpoint for communication with an OS and a user, means for loading a generative AI model, means for analyzing the user's input and generating an appropriate response, means for sending the generated response to a terminal, means for customizing a user interface, means for tracking user operations in real time and displaying additional guidance, and means for analyzing problems, identifying possible causes, and generating solutions. This enables personalized support based on the user's individual needs and operation history, and also enables quick and accurate troubleshooting.
[1443] "Natural language input from a user" refers to a message expressed in everyday language that is input by a user in the form of text, voice, or the like.
[1444] "Means for analyzing natural language input" refers to an algorithm or system that understands the natural language message received from the user and analyzes its grammatical structure and meaning.
[1445] "Means for generating a corresponding response" refers to a system or algorithm for generating an appropriate reply or instruction based on the results of natural language analysis.
[1446] "Means for displaying a response to a user" refers to an interface or device for displaying the generated response so that the user can visually confirm it.
[1447] "Means for opening API endpoints" refers to a system that provides the protocols and interfaces necessary to communicate with information terminals and operating systems.
[1448] "Means for loading a generative AI model" refers to the process or algorithm used to load a generative model into memory or a processor and make it executable.
[1449] "Means for analyzing user input and generating an appropriate response" refers to a process or system that analyzes user input and generates an optimal response based on the results.
[1450] The "means for transmitting the generated response to the terminal" refers to a system for sending the generated response data to the terminal via a network.
[1451] "Means for customizing a user interface" refers to a method or system for changing the design and functionality of an interface based on user settings and operation history.
[1452] "Means for tracking user operations in real time and displaying additional guidance" refers to a system that monitors the operations performed by a user in real time and provides next steps or additional instructions as needed.
[1453] "Means for analyzing problems, identifying possible causes, and generating solutions" refers to the systems and processes used to analyze user-reported problems, identify their causes, and generate appropriate solutions.
[1454] This invention relates to a chatbot system that supports the use of operating systems (OS) for PCs and smartphones. Specifically, it is a system that uses a generative AI model to analyze natural language input from users and generate appropriate responses.
[1455] This system mainly consists of three elements: a server, a terminal, and a user.
[1456] server
[1457] The server hosts the chatbot application and loads the generative AI model (e.g., GPT-4). The server exposes API endpoints for communication with the OS and users, and receives and analyzes user input. The server has the following capabilities:
[1458] A means of receiving natural language input from the user
[1459] A means of parsing natural language input
[1460] A means of generating a corresponding response based on the analysis results
[1461] A means of sending the generated response to the terminal
[1462] A means of analyzing problems, identifying possible causes, and generating solutions
[1463] Specifically, the server performs the following process.
[1464] 1. When receiving input from the user, retrieve the data through the appropriate API endpoint.
[1465] 2. The received input data is analyzed using a generative AI model, which performs tokenization and grammar analysis to understand meaning.
[1466] 3. Based on the analysis results, the server generates an appropriate response, referencing the user profile and past operation history to provide personalized suggestions.
[1467] 4. The generated response is sent to the device through the API endpoint.
[1468] Terminal
[1469] The device runs a dedicated chatbot application built into the operating system. The device is equipped with the following means:
[1470] A way to customize the user interface
[1471] A means to track user actions in real time and display additional guidance
[1472] Specifically, the terminal operates as follows.
[1473] 1. Launch the chatbot application and display a user interface that allows the user to enter a message.
[1474] 2. If there is any input from the user, send it to the server.
[1475] 3. Display the response received from the server in the user interface, providing interactive instructions as needed.
[1476] 4. Track user actions in real time and provide appropriate guidance when information is missing or additional guidance is needed.
[1477] User
[1478] Users access the chatbot application using a computer or smartphone and receive support through natural language dialogue. Users mainly perform the following operations:
[1479] 1. Log in to the application and provide the required authentication information.
[1480] 2. Enter questions or instructions in the chat box and receive responses from the server.
[1481] 3. Follow the displayed instructions and guides to operate the OS.
[1482] Specific examples
[1483] For example, when a user who has just purchased a new PC performs initial setup, the following dialogue may take place:
[1484] User: "I just started up my new PC. What do I do now?"
[1485] Server: "Nice to meet you. Let's go over some basic setup together. First, let's set up the language."
[1486] User: "Please set it to Japanese."
[1487] Server: "The language setting has been changed to Japanese. Next, set the time zone."
[1488] Also, if the user wants to install a new application, the prompt sentence is as follows:
[1489] Prompt for "I want to install a new app":
[1490] Input to the server: "The user wants to install a new app. Please guide them through the installation process."
[1491] Example response: "Find the right application and guide me through the installation process."
[1492] This invention allows users to intuitively and efficiently use OS functions and settings, and also allows them to receive prompt and accurate support when problems occur.
[1493] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1494] Step 1: Initialization and Setup
[1495] server:
[1496] The server hosts the chatbot application, loads the generative AI model (e.g., GPT-4), opens an API endpoint, and prepares to communicate with the OS and the user. Specific operations include the following steps:
[1497] The server reads specific configuration files and database connection information.
[1498] The server loads the generative AI model into memory and makes it executable.
[1499] Device:
[1500] The device launches a dedicated chatbot app and displays a user interface (UI), allowing the user to enter a message. Specifically, the process involves the following steps:
[1501] The device prepares the UI for display, rendering text boxes, submit buttons, etc.
[1502] The terminal checks the connection to the server and whether communication is possible.
[1503] User:
[1504] A user accesses a chatbot application and starts a conversation. The specific operations include the following steps:
[1505] The user checks the application's UI and enters an initial message such as "Hello."
[1506] Step 2: User authentication and profiling
[1507] User:
[1508] A user logs into an application, typically by entering a username and password, and sometimes by using biometric authentication such as facial recognition or fingerprint authentication.
[1509] server:
[1510] The server checks the user's authentication information against the database and retrieves the user profile. Specifically, the following process is performed:
[1511] The server takes the entered username and password and issues an SQL query to the database to verify it.
[1512] If the authentication is successful, the server obtains the user profile (including operation history and setting information) and stores it in memory.
[1513] Device:
[1514] The device customizes the UI based on the user profile information received from the server. Specifically, the following process is performed:
[1515] The device changes the UI style based on the user's language and theme settings.
[1516] Step 3: Natural Language Understanding and Dialogue Generation
[1517] User:
[1518] The user types a question or instruction into the chat box, for example, "I want to install a new app."
[1519] Device:
[1520] The terminal sends the user's input to the server. The specific procedure is as follows:
[1521] The device receives the input as text data and sends an HTTP POST request through the API endpoint.
[1522] server:
[1523] The server analyzes the received user input using a generative AI model. Specifically, the following data processing and calculations are performed:
[1524] The server tokenizes the input data, performs grammatical analysis, and analyzes it to understand its meaning.
[1525] Based on the analysis results, an appropriate response is generated.
[1526] Device:
[1527] The generated response is sent to the terminal and displayed on the user interface. Specifically, the process includes the following steps:
[1528] The terminal acquires the response data received from the server and displays it in the text box.
[1529] Step 4: Implementing functions and assisting operations
[1530] server:
[1531] The server then generates a response that recommends a specific action, such as instructions for installing a new app.
[1532] Device:
[1533] The terminal guides the user according to the procedure from the server. Specifically, the terminal performs the following procedure.
[1534] The device will display instructions such as "Click the install button."
[1535] User:
[1536] The user follows the instructions displayed on the device, for example, "click the install button."
[1537] Step 5: Troubleshooting
[1538] User:
[1539] The user reports a specific problem, for example, "I can't connect to Wi-Fi."
[1540] server:
[1541] The server analyzes the problem, identifies possible causes, and generates solutions. Specifically, the following data processing and calculations are performed:
[1542] The server references the data for troubleshooting and identifies the cause through natural language analysis.
[1543] The server generates solutions and organizes them into guidance.
[1544] Device:
[1545] The generated troubleshooting guidance is displayed in the user interface, specifically by following the steps below.
[1546] The device will display instructions such as "First, please check your volume settings."
[1547] This system allows users to operate the OS and solve problems intuitively and efficiently.
[1548] (Application example 1)
[1549] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1550] Conventional systems make it difficult for users to operate and configure the system in natural language, and they lack sufficient support for security-related settings and troubleshooting. Furthermore, they do not provide personalized suggestions that effectively utilize user profiles and operation histories, making it difficult for users to operate the system intuitively and efficiently.
[1551] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1552] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, and means for generating a corresponding response based on the analysis, thereby enabling the user to perform security settings and other operations in natural language.
[1553] Furthermore, the server includes means for performing specific actions based on the generated responses and means for analyzing problems and generating solutions during troubleshooting, thereby enabling the user to obtain specific troubleshooting and operation guidance.
[1554] Furthermore, by including a means for acquiring and managing a user profile, a means for providing personalized suggestions to the user based on the profile, and a means for adjusting security settings based on the user's operation history and setting information, intuitive and individually optimized operation becomes possible, allowing the user to efficiently manage security settings.
[1555] "Natural language input" is an input format in which a user gives instructions or asks questions to a system in everyday language.
[1556] "Analysis" refers to the process of converting input natural language into a form that a machine can understand and interpreting its meaning.
[1557] A "response" is a reply or instruction from the system that is generated based on the parsed natural language input.
[1558] A "user profile" is a data set that collects information about a user, including individual settings and past operation history.
[1559] "Personalized suggestions" are suggestions that individually provide users with information and guidance that is most suitable for them based on their user profile.
[1560] The "operation history" is a record of a series of operations and setting changes that the user has performed in the past.
[1561] A "specific action" is a specific operation or process that the system actually performs based on a user's instruction.
[1562] "Troubleshooting" is the process of identifying the cause of a problem and providing a solution.
[1563] A "generative AI model" is an artificial intelligence technology that uses machine learning to understand natural language and generate appropriate responses.
[1564] "Security settings" refers to a set of settings and policies that ensure the security of a device or system.
[1565] As an embodiment of the present invention, we will explain a chatbot system in which a server and a terminal cooperate to assist users in security configuration and troubleshooting. This system utilizes a generative AI model to provide appropriate responses to security-related instructions and questions entered by the user in natural language.
[1566] Programs and hardware / software used
[1567] An API (specifically, the OpenAI API) that runs the generative AI model is installed on the server, which analyzes input from users and generates responses. A smartphone is used as the terminal, and a chatbot app is installed on this smartphone. Users interact with the system through this app.
[1568] Natural language input processing begins when a user enters instructions or questions in natural language into the chatbot app. This input is sent to the server as text data, where it is analyzed by the generative AI model. Based on the results of the analysis, the server generates an appropriate response and sends it back to the device.
[1569] User profile management is achieved by applications sending user information to a server, which then stores the user's operation history and settings information and generates personalized suggestions and actions based on this information, which can be useful when adjusting security settings.
[1570] For example, if a user types "check Wi-Fi settings," the device will execute a script based on the generated response to help check the actual Wi-Fi settings, allowing users to easily troubleshoot security settings and resolve issues.
[1571] When troubleshooting, users report their issues in natural language and their input is sent to the server, which uses generative AI models to analyze the cause of the problem and generate a solution as a response. For example, if a user reports "no sound," specific steps such as "check your volume settings" are provided.
[1572] Specific examples
[1573] User: "Check the security status of your device"
[1574] Chatbot: "We're running a security check on your device. Please wait a moment... All security settings are OK."
[1575] Example prompt sentence:
[1576] User input: Check the security status of your device
[1577] Security Chatbot Response: We're going to perform a security check on your device. Please wait... All your security settings are fine.
[1578] In this way, users can easily perform security-related operations in natural language through the chatbot and receive appropriate troubleshooting guidance when problems arise, enabling users to manage their security settings intuitively and efficiently.
[1579] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1580] Step 1:
[1581] The device launches the chatbot app, and the user inputs instructions or questions in natural language. The input text data is sent to the server via an API call. The server receives the text data and prepares it for analysis by the generative AI model.
[1582] Input: User's natural language input (text data)
[1583] Output: Text data sent to the server
[1584] Step 2:
[1585] The server passes the received natural language input to the generative AI model, which begins the analysis process. The generative AI model analyzes the input natural language, understands its meaning, and generates an appropriate response. The analysis uses the OpenAI API.
[1586] Input: Text data sent to the server
[1587] Output: Analysis results from the generative AI model
[1588] Step 3:
[1589] Based on the generated analysis, the server generates a specific response to the user's instruction, such as "check your Wi-Fi settings" or any other troubleshooting steps required.
[1590] Input: Analysis results from generative AI model
[1591] Output: Response data to the user
[1592] Step 4:
[1593] The server sends the generated response data to the terminal, which receives the response and displays it on the user interface within the chatbot app. The user can then view it and take the next action.
[1594] Input: Response data to the user
[1595] Output: Response data sent to the terminal
[1596] Step 5:
[1597] The user takes action based on the generated response. For example, if the instruction is "Check your Wi-Fi settings," the user opens the Wi-Fi settings screen and adjusts the settings. After completing this action, the user again provides feedback to the chatbot.
[1598] Input: Response data displayed to the user
[1599] Output: User's actual actions
[1600] Step 6:
[1601] If the user wants to input additional instructions to resolve the issue, the process repeats from step 1. For example, if the user inputs "Please also check the volume settings," the server will analyze it again and return the appropriate instructions.
[1602] Input: Additional user instructions (natural language input)
[1603] Output: The analysis process is restarted.
[1604] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1605] As an embodiment of the present invention, we provide a chatbot system that supports the use of operating systems (OS) for PCs and smartphones. In particular, we will explain an embodiment that incorporates an engine that recognizes the user's emotions.
[1606] First, the system has the means to receive natural language input from users, parse it, generate and display a response, and, combined with an emotion engine, recognizes the user's emotional state and adapts responses based on that, achieving a high level of personalization.
[1607] Program processing overview
[1608] 1. Initialization and Setup
[1609] The server hosts the chatbot application and loads the generative AI model, a machine learning model for natural language analysis and response generation, and the emotion engine, a model for recognizing user emotions.
[1610] The chatbot application is installed on the device and launched. The user interface (UI) is displayed, allowing the user to enter a message. The initial screen displays, "Hello. Is there anything I can help you with?"
[1611] The user accesses the chatbot application and starts a dialogue, receiving the user's natural language or voice input in text format.
[1612] 2. User Authentication and Profiling
[1613] A user logs in to an application. The server checks the user information and retrieves the user profile. The retrieved user profile includes past operation history and setting information.
[1614] The device displays a personalized UI based on the user's settings and past operation history, allowing it to provide more appropriate suggestions and guidance to the user.
[1615] 3. Natural Language Understanding and Emotion Recognition
[1616] Users enter questions or instructions into the chat box (e.g., "I want to install a new app"). The system also analyzes the user's facial expressions and tone of voice using an emotion engine. The device then sends these inputs to the server.
[1617] The server uses a generative AI model to analyze natural language and an emotion engine to recognize the user's emotions. The analysis results identify the user's intent, request, and emotional state (e.g., "I want to know the steps to install a new app. The user is feeling a bit anxious and worried.").
[1618] 4. Response generation and emotional adaptation
[1619] Based on the analysis and emotional state, the server generates an appropriate response, such as a command to open the application store or friendly, gentle installation instructions.
[1620] The generated response is sent from the server to the device, and may include specific instructions or an emotionally sensitive tone.
[1621] The device displays the received response on the user interface, for example, a message such as "Open the application store and install the desired app. If you have any questions, please feel free to ask."
[1622] 5. Troubleshooting
[1623] When a user reports a specific problem (e.g., "I can't connect to Wi-Fi"), the server analyzes the problem and evaluates the user's emotional state using an emotion engine. The generated troubleshooting guidance is sent to the device as specific solution steps. The device then displays the solution steps on the user interface and provides emotion-sensitive instructions.
[1624] Specific examples
[1625] New user onboarding
[1626] User: Starts up a newly purchased PC and uses the chatbot for the first time.
[1627] Device: Display "Hello. Let's go through the basic setup together. If you have any questions, please feel free to ask."
[1628] User: Type "Please set to Japanese."
[1629] Terminal: Send this input to the server.
[1630] Server: Parses the message and generates instructions to change the language setting to Japanese. Recognizes that the user is feeling a bit anxious and generates a response in a friendly tone.
[1631] Device: Change the language setting to Japanese and notify the user, "Setup complete. You can now use Japanese. Is there anything else we can help you with?"
[1632] Introducing and explaining new features
[1633] User: After an update, launch the chatbot and type, "Tell me about the new features."
[1634] Server: Searches for information about new features and generates summaries of them, recognizing the user's emotional state and crafting interesting descriptions.
[1635] On your device: "Our new update features a fresh design, battery optimizations, and new security features."
[1636] Fault diagnosis and troubleshooting
[1637] User: Reports "No sound."
[1638] Server: Analyzes the problem and recognizes the user's emotional state using an emotion engine. Lists possible causes and generates solutions.
[1639] Device: "Possible reasons for the lack of sound include the volume setting, muting, or a faulty speaker. First, check the volume setting. If you have any questions, please let us know."
[1640] User: Follow the instructions and check the volume settings again. In this way, a system incorporating an emotion engine allows users to receive responses that take their emotions into consideration, allowing them to operate the OS with greater peace of mind.
[1641] The processing flow will be explained below.
[1642] Step 1:
[1643] The server launches the chatbot application and loads the generative AI model and emotion engine. The generative AI model is a machine learning model for natural language analysis and response generation, and the emotion engine is a model for recognizing user emotions. The server also opens an API endpoint and prepares to accept connection requests from devices.
[1644] Step 2:
[1645] The chatbot application is installed on the device and launched. The user interface (UI) is displayed, allowing the user to enter a message. The initial screen displays, "Hello. Is there anything I can help you with?"
[1646] Step 3:
[1647] The user enters login information (e.g., user ID and password) into the UI and clicks the "Login" button, which sends a login request from the terminal to the server.
[1648] Step 4:
[1649] The server compares the received login information with a database and performs authentication. If authentication is successful, the server obtains the user profile and sends it to the terminal. The user profile includes past operation history and individual setting information.
[1650] Step 5:
[1651] The device will then display a personalized UI based on the received user profile, such as shortcuts to frequently used functions and recently accessed applications.
[1652] Step 6:
[1653] Users enter questions or instructions into the chat box (e.g., "I want to install a new app"). The entered message is sent directly to the server. The system also analyzes the user's facial expressions and tone of voice using an emotion engine.
[1654] Step 7:
[1655] The server receives the user's input message and uses a generative AI model to perform natural language analysis, identifying the user's intent and request (e.g., "I want to know the steps to install a new app").
[1656] Step 8:
[1657] The server uses an emotion engine to recognize the user's emotional state, for example, to determine whether the user is feeling anxious or worried.
[1658] Step 9:
[1659] The server generates an appropriate response based on the analysis and emotional state, for example, "Do you want to open the app store?" in a calm and reassuring tone.
[1660] Step 10:
[1661] The generated response is sent from the server to the device, and may include specific instructions or an emotionally sensitive tone.
[1662] Step 11:
[1663] The device displays the received response on the user interface, for example, "Open the application store and install the desired app. If you have any questions, please feel free to ask."
[1664] Step 12:
[1665] The user follows the displayed instructions (e.g., answers "yes" and opens the application store), and the user's selection is sent back to the server.
[1666] Step 13:
[1667] Based on the user's selection, the server determines the next action and sends the necessary instructions to the terminal, such as to display search results for a particular application.
[1668] Step 14:
[1669] The device updates its user interface based on the instructions received from the server, for example by displaying a list of search results and allowing the user to select the app they want to install.
[1670] Step 15:
[1671] The user selects the app they want to install and sends an installation request to the server again.
[1672] Step 16:
[1673] The server confirms the installation request, initiates the necessary installation process, and notifies the device of the installation progress.
[1674] Step 17:
[1675] The terminal displays the installation progress to the user and notifies the user when the installation is complete.
[1676] Step 18:
[1677] Once the installation is complete, the user launches the application and begins using it. Through this process, the user can interact with the OS based on natural language and emotions to smoothly perform the required actions.
[1678] Example 2
[1679] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1680] Conventional chatbot systems have difficulty generating appropriate responses to users' natural language input, and in particular, they are unable to recognize emotions and reflect them in responses, which means they are unable to provide a satisfying conversational experience for users.In addition, they lack personalized suggestions that utilize the user's profile information and past operation history, making it difficult to achieve user-friendly operation.
[1681] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1682] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, means for generating a corresponding response based on the analysis, means for displaying the response to the user, means for analyzing natural language using a generative AI model, means for identifying the user's emotional state using an emotion recognition engine, means for generating an appropriate response based on the analysis result and the emotional state, and means for displaying the response in a tone that takes the user's emotions into consideration. This makes it possible to provide more personalized suggestions and guidance while taking the user's emotions into consideration, thereby achieving a highly satisfying dialogue experience.
[1683] "Natural language input" refers to text or voice input entered by a user in the form of everyday spoken language.
[1684] A "generative AI model" refers to an algorithmic model that uses machine learning technology to analyze natural language and generate responses.
[1685] "Emotion recognition engine" refers to a software module for analyzing and identifying an emotional state from a user's input.
[1686] A "user profile" refers to a collection of data including a user's personal information, past operation history, and setting information.
[1687] "Personalization" refers to providing customized suggestions and responses based on individual user characteristics and past behavior.
[1688] "Operation history" refers to a record of operations that a user has performed in the past.
[1689] "Setting information" refers to information about various settings that a user makes when using an application or system.
[1690] A "server" refers to a computer or software system used to receive requests from users and process or respond to data.
[1691] "Terminal" refers to the device that the user actually operates (e.g., a PC or smartphone).
[1692] "User interface" refers to the screen display and input means that allow a user to interact with a system.
[1693] "Login" refers to the authentication procedure required for a user to access a system.
[1694] "Analysis" refers to the processing of data to understand a user's natural language input and emotional state.
[1695] A "response" refers to a reply or instruction that a system generates and provides in response to a user's input.
[1696] As an embodiment of the present invention, we provide a chatbot system that supports the use of operating systems (OS) for personal computers and smartphones. This system incorporates a generative AI model and an emotion recognition engine.
[1697] First, the server hosts the chatbot application and loads a generative AI model (e.g., GPT-3.5) and an emotion recognition engine (e.g., a popular emotion recognition API). The generative AI model is used to parse natural language and generate responses, and the emotion recognition engine is used to recognize the user's emotional state. These models and engines are initialized by the server by loading the API key and model data.
[1698] The device then allows the user to install and launch the chatbot application. When the user launches the application, a user interface appears, displaying a welcome message that reads, "Hello, how can I help you?"
[1699] The user accesses the chatbot application and starts a dialogue using natural language or voice. The input content is received in text format and sent from the terminal to the server. If the user inputs voice, the voice is converted into text and sent to the server.
[1700] When a user logs in to an application, the server checks the user information and obtains profile data (past operation history, setting information, etc.). Based on the obtained profile data, a personalized UI is displayed to the user. For example, frequently used functions can be displayed as shortcuts based on the past operation history.
[1701] The server analyzes the natural language input from the user using a generative AI model to identify the user's intentions and requests. It also uses an emotion recognition engine to recognize the user's emotional state. For example, if a user inputs "I want to install a new app," the analysis results might determine that "I want to know the steps to install a new app, and I'm feeling a bit impatient and anxious."
[1702] Based on the analysis results and the emotional state, the server uses a generative AI model to generate an appropriate response, such as "Please visit the application store and install the app you want. If you have any questions, please feel free to ask." This response is generated in an emotionally sensitive tone.
[1703] The generated response is sent to the device and displayed in the user interface. The displayed response is provided in an emotionally sensitive tone, allowing the user to proceed with the operation with confidence.
[1704] When a user reports a specific problem, for example, "I can't connect to Wi-Fi," the server analyzes the problem and uses an emotion recognition engine to assess the user's emotional state. Based on the analysis, specific solutions are generated and sent to the device. For example, a message might read, "Possible causes of the lack of audio include volume settings, muting, or a faulty speaker. First, check your volume settings."
[1705] Specific examples
[1706] New user onboarding
[1707] User: Starts up a newly purchased PC and uses the chatbot for the first time.
[1708] Device: Display "Hello. Let's go through the basic setup together. If you have any questions, please feel free to ask."
[1709] User: Type "Please set to Japanese."
[1710] Terminal: Send this input to the server.
[1711] Server: Parses the message and generates instructions to change the language setting to Japanese. Recognizes that the user is feeling a bit anxious and generates a response in a friendly tone.
[1712] Device: Change the language setting to Japanese and notify the user, "Setup complete. You can now use Japanese. Is there anything else we can help you with?"
[1713] Introducing and explaining new features
[1714] User: After an update, launch the chatbot and type, "Tell me about the new features."
[1715] Server: Searches for information about new features and generates summaries of them, recognizing the user's emotional state and crafting interesting descriptions.
[1716] On your device: "Our new update features a fresh design, battery optimizations, and new security features."
[1717] Fault diagnosis and troubleshooting
[1718] User: Reports "No sound."
[1719] Server: Analyzes the problem and recognizes the user's emotional state using an emotion recognition engine. Lists possible causes and generates solutions.
[1720] Device: "Possible reasons for the lack of sound include the volume setting, muting, or a faulty speaker. First, check the volume setting. If you have any questions, please let us know."
[1721] In this way, a system incorporating an emotion recognition engine allows users to receive responses that take their emotions into consideration, allowing them to operate the OS with greater peace of mind.
[1722] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1723] Step 1:
[1724] The server hosts the chatbot application and loads the generative AI model (e.g., GPT-3.5) and emotion recognition engine (e.g., a popular emotion recognition API). Specifically, it loads the API key and model data to initialize the model. The input at this stage is the required API key and model data, and the output is a ready-to-use generative AI model and emotion recognition engine.
[1725] Step 2:
[1726] When a user installs and launches a chatbot application, the device loads the user interface. The initial screen displays "Hello, is there anything I can help you with?" When a user launches the application, the input is the application launch command, and the output is a user interface displaying the message.
[1727] Step 3:
[1728] The user inputs questions or instructions to the chatbot using natural language or voice. For example, they might input "I want to install a new app." This input is received by the device in text format. If the input is voice, the voice is converted to text. At this stage, the input is the user's natural language or voice input, and the output is text data.
[1729] Step 4:
[1730] The device sends the received text data to the server. It also sends text data converted from speech. Specifically, it sends the text data to the server via an HTTP request. At this stage, the input is user data in text format, and the output is the data sent to the server.
[1731] Step 5:
[1732] The server analyzes the received text data using a generative AI model to identify the user's intentions and requests. It also uses an emotion recognition engine to recognize the user's emotional state. For example, if a user inputs "I want to install a new app," the analysis results might determine that "I want to know the steps to install a new app. The user is feeling a bit impatient and anxious." The input at this stage is text data, and the output is the analysis results regarding the user's intentions and emotional state.
[1733] Step 6:
[1734] The server uses a generative AI model to generate an appropriate response based on the analysis results and the emotional state. For example, it generates a message such as, "Open the application store and install the app you want. If you have any questions, please feel free to ask." This response is generated in an emotionally sensitive tone. The input at this stage is the analysis results and the emotional state data, and the output is an appropriate response message.
[1735] Step 7:
[1736] The server sends the generated response message to the device. The device receives the response and displays it on its user interface. The displayed message contains specific instructions such as "Open the application store and install the desired app." The input at this stage is the response message, and the output is the message displayed on the user interface.
[1737] Step 8:
[1738] When a user reports a specific problem, for example, "I can't connect to Wi-Fi," the server analyzes the problem and uses an emotion recognition engine to assess the user's emotional state. The analysis results in a list of possible causes and generates a solution. For example, a message might read, "Possible causes of the no sound include the volume setting, muting, or a faulty speaker. Please check the volume setting first." The input at this stage is text data about the specific problem, and the output is the problem analysis and a solution message.
[1739] Step 9:
[1740] The terminal displays the solution message received from the server on the user interface. For example, it displays "Please check your volume settings." This message is expressed in an emotionally sensitive tone. The input at this stage is the solution message, and the output is the solution message displayed on the user interface.
[1741] This allows users to receive emotionally sensitive responses and operate with peace of mind.The system improves the user experience by combining natural language input analysis and emotion recognition.
[1742] (Application example 2)
[1743] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1744] In factory work environments, workers are required to be able to operate robots and troubleshoot them quickly and reliably. However, conventional systems have difficulty responding to the user's emotional state and real-time work environment, which often causes stress and confusion for workers. This can hinder efficient work execution and reduce safety.
[1745] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1746] In this invention, the server includes means for receiving natural language input from a user, means for analyzing the natural language input, means for generating a corresponding response, means for recognizing the user's emotional state, and means for adapting the response based on the user's emotional state. This enables generation of a response that takes the user's emotions into consideration. The server also includes means for acquiring the user's operation history and setting information, means for providing information related to the work environment in real time, and means for providing appropriate operation guidance to the user. This enables appropriate guidance based on real-time information about the work environment, allowing work to be performed safely and efficiently.
[1747] "Natural language input" refers to a language that is input by a user in the form of text, speech, or other language that is commonly used by humans.
[1748] "Analysis" refers to the syntactic and semantic processing and interpretation of input natural language.
[1749] A "response" is a response to a question or instruction that is generated based on the results of the analysis.
[1750] "Display" refers to providing the generated response to the user visually or audibly.
[1751] "Emotional state" refers to the user's psychological and emotional state as determined by their tone of voice and facial expression.
[1752] A "profile" is a collection of data about a user, including the user's personal information, operation history, setting information, and the like.
[1753] "Personalization" refers to providing individualized offers and information based on a user's profile information.
[1754] An "operation history" is a record of operations and actions that a user has performed in the past.
[1755] "Setting information" refers to setting items and parameters that are individually set by the user for the system or application.
[1756] "Work environment information" is information relating to the user's real-time work location and situation.
[1757] "Operational guidance" refers to procedures or instructions provided to a user to perform a specific task or operation.
[1758] This embodiment of the present invention is a chatbot system with advanced response capabilities incorporating user emotion recognition. It is particularly intended for supporting robotic operation in factory work environments. The system analyzes the user's natural language input, generates corresponding responses, and displays those responses. It also provides appropriate operational guidance in real time based on the user profile and work environment information.
[1759] 1. Hardware and Software Used
[1760] This system uses the following hardware and software:
[1761] Hardware: Smartphones, smart glasses, head-mounted displays
[1762] software:
[1763] Python: To execute the main logic of the program
[1764] OpenAI API: Natural Language Analysis and Response Generation
[1765] Emotion Recognition API: To recognize user emotions
[1766] 2. Program processing content
[1767] The server uses a generative AI model to analyze natural language and generate appropriate responses to user input. Furthermore, it uses the Emotion Recognition API to determine the user's emotional state and adapts responses based on that emotion. Specific responses include:
[1768] 3. System Operation Example
[1769] Operation Guidance
[1770] When a user inquires about how to operate the robot, the server analyzes the information and provides appropriate operating instructions. For example, if a user asks, "How do I install a new part?", the server generates a response like this: "The installation steps are as follows: First, turn off the power and check the safety devices. Then install the part. Please let us know if you have any questions."
[1771] troubleshooting
[1772] When a user reports a problem with the robot, the server analyzes the cause and provides instructions based on emotion recognition to reduce stress and confusion. For example, if a user types, "My robot has stopped working, what should I do?", the server generates a response such as, "Please check the robot's power status and try restarting it. Is there anything else I can help you with?"
[1773] 4. Examples of prompts
[1774] "My robot has stopped working, what should I do?"
[1775] "How do I install the new parts?"
[1776] This allows the user to receive accurate work guidance in real time, improving the efficiency and safety of robot operation, and serves as a clear solution to the problem that the invention aims to solve.
[1777] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1778] Step 1:
[1779] The server hosts the chatbot application and loads the generative AI model and emotion recognition engine. The device installs and launches the chatbot application. The user interface is displayed, allowing the user to enter a message. When the user launches the application, the initial screen displays, "Hello. Is there anything I can help you with? How can I help you operate your robot?"
[1780] Step 2:
[1781] The user logs in to the application. The server checks the user information and retrieves the user profile. The retrieved user profile includes past operation history and setting information. The device displays a personalized UI based on the user's settings and past operation history. This allows the device to provide the user with appropriate suggestions and guidance.
[1782] Step 3:
[1783] The user enters questions or instructions into the chat box. For example, "How do I install a new part?" The system also analyzes the user's voice and facial expressions using an emotion engine. The device then sends these inputs to the server. The entered text data and voice data are then sent to the server.
[1784] Step 4:
[1785] The server uses a generative AI model to analyze natural language and an emotion recognition engine to recognize the user's emotional state. The analysis results identify the user's intentions, requests, and emotional state. For example, it may determine that the user wants to know the procedure for installing a new part. The user is feeling a little anxious. Natural language analysis extracts keywords from the input text data, and emotion recognition is performed by analyzing voice tone and facial expression data.
[1786] Step 5:
[1787] The server generates an appropriate response based on the analysis results and the user's emotional state. For example, it may describe the installation procedure in detail in a calm tone. It may generate a response such as, "The installation procedure is as follows: First, turn off the power and check the safety devices. Then, install the parts. Please let us know if you have any questions." The generated response is in text format.
[1788] Step 6:
[1789] The generated response is sent from the server to the terminal. The terminal displays the received response on the user interface. Based on the displayed message, the user can obtain specific instructions for performing the next action. For example, a message such as "The installation procedure is as follows. First, turn off the power and check the safety devices. Then install the parts. Please let us know if you have any questions" may be displayed.
[1790] Step 7:
[1791] When a user reports a specific problem, for example, "My robot has stopped working, what should I do?", the server analyzes the problem and evaluates the user's emotional state using an emotion recognition engine. The generated troubleshooting guidance is sent to the device as specific solution steps. The device displays the solution steps on the user interface and provides emotion-sensitive instructions. For example, a message such as "Please check the robot's power status and try restarting it. Is there anything else I can help you with?" is displayed.
[1792] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1793] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1794] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1795] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1796] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1797] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1798] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1799] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1800] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1801] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1802] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1803] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1804] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1805] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1806] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1807] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1808] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1809] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1810] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1811] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1812] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1813] The following is further disclosed regarding the above embodiment.
[1814] (Claim 1)
[1815] means for receiving natural language input from a user;
[1816] means for analyzing said natural language input;
[1817] means for generating a corresponding response based on said analysis;
[1818] The system includes means for displaying said response to a user.
[1819] (Claim 2)
[1820] a means for obtaining and managing user profiles;
[1821] 10. The system of claim 1, further comprising means for providing personalized suggestions to the user based on the profile.
[1822] (Claim 3)
[1823] A means for acquiring user operation history and setting information;
[1824] 2. The system according to claim 1, further comprising means for assisting a user in searching for an application or setting desired by the user using the operation history and setting information.
[1825] (Claim 4)
[1826] means for receiving fault reports from users and diagnosing problems;
[1827] 10. The system of claim 1, further comprising means for generating troubleshooting guidance based on the diagnostic results.
[1828] (Claim 5)
[1829] A means of obtaining and notifying users of new features and updates for the OS;
[1830] The system of claim 1 further comprising means for explaining how to use the new features and updates.
[1831] (Claim 6)
[1832] A means of providing voice and text interaction for users with visual limitations or language differences;
[1833] 10. The system of claim 1, further comprising means for enhancing interaction with the OS through said dialogue.
[1834] (Claim 7)
[1835] A means of efficiently managing the simultaneous operation of multiple applications and functions;
[1836] The system according to claim 1, further comprising means for supporting switching and cooperation of the simultaneous operations.
[1837] "Example 1"
[1838] (Claim 1)
[1839] means for receiving natural language input from a user;
[1840] means for analyzing said natural language input;
[1841] means for generating a corresponding response based on said analysis;
[1842] means for displaying said response to a user;
[1843] A means to open API endpoints for communication with the OS and users,
[1844] a means for loading the generative AI model;
[1845] means for parsing user input and generating an appropriate response;
[1846] means for transmitting the generated response to the terminal;
[1847] a means for customizing the user interface;
[1848] means for tracking user actions in real time and displaying additional guidance;
[1849] A means of analyzing problems, identifying possible causes, and generating solutions
[1850] A system including:
[1851] (Claim 2)
[1852] a means for obtaining and managing user profiles;
[1853] 10. The system of claim 1, further comprising means for providing personalized suggestions to the user based on the profile.
[1854] (Claim 3)
[1855] A means for acquiring user operation history and setting information;
[1856] 2. The system according to claim 1, further comprising means for assisting a user in searching for an application or setting desired by the user using the operation history and setting information.
[1857] "Application Example 1"
[1858] (Claim 1)
[1859] means for receiving natural language input from a user;
[1860] means for analyzing said natural language input;
[1861] means for generating a corresponding response based on said analysis;
[1862] means for displaying said response to a user;
[1863] a means for taking specific actions based on the generated responses;
[1864] A means of analyzing problems and generating solutions when troubleshooting;
[1865] A system including:
[1866] (Claim 2)
[1867] a means for obtaining and managing user profiles;
[1868] means for providing personalized suggestions to the user based on said profile;
[1869] 2. The system according to claim 1, further comprising means for adjusting security settings based on a user's operation history and setting information.
[1870] (Claim 3)
[1871] A means for acquiring user operation history and setting information;
[1872] a means for assisting a user in searching for applications and settings desired by the user using the operation history and setting information;
[1873] 10. The system of claim 1, further comprising means for using a generative AI model to generate appropriate responses from natural language inputs to troubleshoot security-related issues.
[1874] "Example 2: Combining Emotion Engines"
[1875] (Claim 1)
[1876] means for receiving natural language input from a user;
[1877] means for analyzing said natural language input;
[1878] means for generating a corresponding response based on said analysis;
[1879] means for displaying said response to a user;
[1880] A means of analyzing natural language using generative AI models; and
[1881] means for identifying an emotional state of a user utilizing an emotion recognition engine;
[1882] means for generating an appropriate response based on the analysis and the emotional state;
[1883] means for displaying said response in an emotionally sensitive tone.
[1884] (Claim 2)
[1885] a means for obtaining and managing user profiles;
[1886] 10. The system of claim 1, further comprising means for providing personalized suggestions to the user based on the profile.
[1887] (Claim 3)
[1888] A means for acquiring user operation history and setting information;
[1889] 2. The system according to claim 1, further comprising means for assisting a user in searching for a desired function or setting by using the operation history and setting information.
[1890] "Application example 2 when combining emotion engines"
[1891] (Claim 1)
[1892] means for receiving natural language input from a user;
[1893] means for analyzing said natural language input;
[1894] means for generating a corresponding response based on said analysis;
[1895] means for displaying said response to a user;
[1896] means for recognizing the emotional state of a user;
[1897] means for adapting a response based on said emotional state;
[1898] Including system.
[1899] (Claim 2)
[1900] a means for obtaining and managing user profiles;
[1901] 10. The system of claim 1, further comprising means for providing personalized suggestions to the user based on the profile.
[1902] (Claim 3)
[1903] A means for acquiring user operation history and setting information;
[1904] a means for assisting a user in searching for applications and settings desired by the user using the operation history and setting information;
[1905] means for providing real-time information related to a user's work environment;
[1906] means for providing appropriate operation guidance to a user based on the work environment information;
[1907] 10. The system of claim 1, comprising: [Explanation of symbols]
[1908] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving natural language input from a user; means for analyzing said natural language input; means for generating a corresponding response based on said analysis; The system includes means for displaying said response to a user.
2. a means for obtaining and managing user profiles; The system of claim 1 further comprising means for providing personalized suggestions to the user based on the profile.
3. A means for acquiring user operation history and setting information; 2. The system according to claim 1, further comprising means for assisting a user in searching for an application or setting desired by the user using the operation history and setting information.
4. means for receiving fault reports from users and diagnosing problems; The system of claim 1 further comprising means for generating troubleshooting guidance based on the diagnostic results.
5. A means of obtaining and notifying users of new features and updates for the OS; The system of claim 1 further comprising means for explaining how to use the new features and updates.
6. A means of providing voice and text interaction for users with visual limitations or language differences; The system of claim 1 further comprising means for enhancing interaction with the OS through said dialogue.
7. A means of efficiently managing the simultaneous operation of multiple applications and functions; The system according to claim 1 , further comprising means for supporting switching and cooperation of the simultaneous operations.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A