system

A system that processes natural language inputs to generate and present scripts or commands efficiently, addressing the inefficiencies of manual searches and technical knowledge requirements in conventional systems.

JP2026062156APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Conventional systems require users to refer to manuals or search the internet for information, which is time-consuming and inefficient, especially for those lacking technical knowledge, making it difficult to create appropriate scripts or commands.

Method used

A system that allows users to input instructions in natural language, which are analyzed using a natural language processing engine, generating and presenting appropriate scripts or commands quickly and easily.

Benefits of technology

Enables users to generate and execute scripts and commands efficiently without specialized knowledge, supporting quick and accurate task completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062156000001_ABST
    Figure 2026062156000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] An input means in which the user inputs in natural language, An analysis means for analyzing user input, A generation means that generates a script or command based on the analyzed content, A means of presenting the generated script or command to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a conventional system, a user has to refer to a standard manual or search for information on the Internet, and these methods have problems of taking time and effort. Also, for users lacking technical knowledge, it is difficult to create appropriate scripts or commands. As a result, the efficiency of operations has often been hindered. The present invention aims to solve these problems by providing a system that generates and presents appropriate scripts or commands simply by a user inputting an instruction in natural language.

Means for Solving the Problems

[0005] The present invention solves the above problems by the following means. The system of the present invention

[0006] 1. An input method in which the user inputs in natural language,

[0007] 2. An analysis means for analyzing the user's input,

[0008] 3. A generation means for generating a script or command based on the analyzed content,

[0009] 4. A means for presenting the generated script or command to the user,

[0010] This includes the following: The analysis means uses a natural language processing engine, and the generation means sends the generated script or command to the server, enabling the user to receive the appropriate script or command quickly. This allows the user to proceed with their work quickly and easily.

[0011] A "user" is a person or organization that uses the system to input instructions in natural language and generate and execute scripts and commands.

[0012] "Natural language" refers to the language that humans use in everyday life, not programming languages, but a form of expression that is easily understood by humans.

[0013] "Input method" refers to an interface for users to input instructions in natural language, and includes keyboards, mice, touchscreens, etc.

[0014] "Analysis means" refers to software or hardware that has the function of analyzing the user's natural language input and understanding its content.

[0015] "Generation means" refers to software or hardware that has the function of automatically creating appropriate scripts or commands based on the analyzed user instructions.

[0016] The "prompting means" is an interface for visually displaying a script or command created by the generating means to the user.

[0017] A "script" is a program that describes a series of instructions for automatically executing certain processes.

[0018] A "command" is an instruction sentence for instructing a specific operation to a system or software.

[0019] A "natural language analysis engine" is a technology that analyzes text written in natural language, understands its meaning, and leads to appropriate actions.

[0020] A "server" is a computer system that provides services to other computers (clients) on a network.

Brief Description of Drawings

[0021] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8]A conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] Shows an emotion map to which multiple emotions are mapped. [Figure 10] Shows an emotion map to which multiple emotions are mapped. [Figure 11] A sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] A sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] A sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] A sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0022] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0023] First, the terms used in the following description will be explained.

[0024] In the following embodiments, a processor with a reference number (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0025] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0026] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0027] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0029] [First Embodiment]

[0030] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0031] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0032] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0033] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0034] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0036] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0037] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0038] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0039] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0040] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0041] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0042] Overall Overview

[0043] The present invention provides a system that automatically generates and presents necessary scripts and commands simply by having the user input instructions in natural language. This system supports users in carrying out their tasks quickly and accurately.

[0044] System Configuration

[0045] This system includes the following main components:

[0046] 1. Input method: The user inputs using natural language.

[0047] 2. Analysis method: Analyzes the user's input.

[0048] 3. Generation method: Generate a script or command based on the analyzed content.

[0049] 4. Presentation method: The generated script or command is presented to the user.

[0050] Details of the example

[0051] 1. User input

[0052] The user inputs instructions into the terminal using natural language. For example, they might input, "I want to generate a command to connect to a specified server and check disk usage."

[0053] 2. Natural language analysis

[0054] The terminal sends user input to the server. The server analyzes the input using a natural language processing engine. The processing engine understands what the user wants to achieve and extracts information to generate appropriate scripts or commands.

[0055] 3. Script / command generation

[0056] Based on the analyzed information, the server generates appropriate scripts and commands. For example, the server generates the Unix command "df -h" to check disk usage.

[0057] 4. Presentation of results

[0058] The server generates scripts and commands, which are then sent back to the terminal. The terminal then presents these to the user. The user can review the generated commands and use them.

[0059] Specific example

[0060] Example 1: Checking disk usage

[0061] User input: The user instructs the terminal to "generate a command to check the server's disk usage."

[0062] Server analysis and generation: The server analyzes the user's instructions and generates the disk usage command "df -h".

[0063] Prompt: The terminal prompts the user to execute "df -h". The user uses the suggested command to check the server's disk usage.

[0064] Example 2: Retrieving a file list

[0065] User input: The user instructs the terminal to "generate a command to retrieve a list of files in the specified directory."

[0066] Server analysis and generation: The server analyzes the data using commands like "ls -la / path / to / directory" and generates the appropriate commands.

[0067] Presentation: The terminal presents the user with the command "ls -la / path / to / directory", and the user executes the command to obtain a list of files.

[0068] Operation flow and effects

[0069] This system allows users to quickly and easily generate appropriate scripts and commands, enabling them to work efficiently. Because users can generate complex commands simply by inputting instructions in natural language, the system can be used without problems even by those lacking technical knowledge. Furthermore, the generated scripts and commands can be visually reviewed, allowing users to proceed with confidence.

[0070] The following describes the processing flow.

[0071] Step 1:

[0072] The user enters instructions into the terminal using natural language. For example, they might enter, "I want to generate a command to connect to a specified server and check disk usage."

[0073] Step 2:

[0074] The terminal receives user input and sends that input to the server. The data sent to the server is natural language text data.

[0075] Step 3:

[0076] The server sends the data to a natural language processing engine to analyze the user's input. The natural language processing engine analyzes the input sentence and extracts information to understand the user's intent.

[0077] Step 4:

[0078] The natural language processing engine sends the analysis results back to the server. The analysis results include the user's desired objective (for example, "check disk usage").

[0079] Step 5:

[0080] The server generates an appropriate script or command based on the analysis results. For example, it generates the command "df -h" to check disk usage.

[0081] Step 6:

[0082] The server sends the generated script or command back to the terminal. The terminal receives this data.

[0083] Step 7:

[0084] The terminal displays the received script or command to the user. The user confirms the displayed content on the terminal screen.

[0085] Step 8:

[0086] If a user needs to execute a presented script or command, they send the command to the server via the terminal. The server executes the received command and returns the result to the user.

[0087] Step 9:

[0088] The terminal displays the execution results of commands received from the server to the user. The user can then obtain the necessary information.

[0089] Through the above processing steps, users can efficiently perform a series of operations, from natural language input to the generation and execution of appropriate scripts and commands. This system makes it easy to generate and execute complex commands, even for those lacking technical knowledge.

[0090] (Example 1)

[0091] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0092] In conventional systems, users needed knowledge of commands and scripts to perform the desired operations, making them difficult for non-specialized users. Furthermore, generating scripts and commands had to be done manually, which was time-consuming and laborious. Therefore, there was a need for a system that would allow users to easily generate commands using natural language.

[0093] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0094] In this invention, the server includes an input means for the user to input in natural language, a transmission means for the terminal to send the user's input to the server, an analysis means for the server to analyze the user's input, a generation means to generate a script or command based on the analyzed content, and a presentation means to return the script or command generated from the server to the terminal and present it to the user. This makes it possible for the user to easily generate the desired script or command in natural language and to execute operations quickly and accurately.

[0095] "Natural language" refers to the language that humans use on a daily basis, and is not a specific programming language or specialized terminology, but rather a general language.

[0096] "Input method" refers to a device or interface that allows a user to input information or instructions in natural language.

[0097] "Transmission means" refers to the communication functions and protocols that a terminal uses to send user input to a server.

[0098] "Analysis means" refers to algorithms and engines that semantically understand the natural language input received by the server and perform appropriate processing.

[0099] A "natural language processing engine" refers to software or a system that analyzes input natural language and understands its meaning and intent.

[0100] "Generation means" refers to algorithms or engines used to generate specific scripts or commands based on the analyzed content.

[0101] A "template database" refers to a set of predefined templates that a generation tool references when generating scripts or commands.

[0102] "Presentation means" refers to devices or interfaces used to display generated scripts or commands to the user.

[0103] A "terminal" refers to a computer or mobile device used by a user to input information and receive generated scripts or commands.

[0104] A "server" refers to a computer system equipped with analysis and generation capabilities, designed to process input from terminals.

[0105] Overall Overview

[0106] The present invention provides a system that automatically generates and presents necessary scripts and commands simply by having the user input instructions in natural language. This allows users to perform operations efficiently even without specialized knowledge. The system provides support for users to carry out their tasks quickly and accurately.

[0107] System Configuration

[0108] This system includes the following main components:

[0109] 1. Input method: The user inputs using natural language.

[0110] 2. Transmission method: The terminal sends the user's input to the server.

[0111] 3. Analysis method: The server analyzes the user's input.

[0112] 4. Generation method: Generate a script or command based on the analyzed content.

[0113] 5. Presentation method: The generated script or command is presented to the user.

[0114] Details of the example

[0115] 1. User input

[0116] The user inputs instructions into the terminal using natural language. For example, they might input, "Please generate a command to connect to the specified server and check disk usage." An example of a specific prompt is as follows:

[0117] Example of a prompt:

[0118] "Please generate a command to connect to the specified server and check disk usage."

[0119] "Generate a command to retrieve a list of files in a specified directory."

[0120] 2. Natural language analysis

[0121] The terminal sends the user's input to the server. The server analyzes the input using a natural language processing engine (e.g., "GPT-4®"). The natural language processing engine syntactically analyzes the input text, understands its intent, and then extracts the necessary information.

[0122] 3. Script / command generation

[0123] Based on the analyzed data, the server generates appropriate scripts and commands. In this process, the server uses a command generation engine (e.g., a "shell script generation library") to generate specific commands by supplementing templates obtained from a template database. For example, to check disk usage, the Unix command "df -h" is generated.

[0124] 4. Presentation of results

[0125] The generated scripts and commands are sent back from the server to the terminal. The terminal visually presents them to the user, who can then review and use them. For example, the generated command "df -h" is displayed on the terminal screen and executed by the user.

[0126] Specific example

[0127] Example 1: Checking disk usage

[0128] User input: "Generate a command to check the server's disk usage."

[0129] Server analysis and generation: The server analyzes the input and generates the command "df -h".

[0130] Instructions: The terminal will display "df -h" to the user, who will use this to check disk usage.

[0131] Example 2: Retrieving a file list

[0132] User input: "Generate a command to get a list of files in the specified directory."

[0133] Server analysis and generation: The server generates the command "ls -la / path / to / directory".

[0134] Instructions: The terminal displays "ls -la / path / to / directory" to the user, who then uses this to obtain a list of files.

[0135] Operation flow and effects

[0136] This system allows users to quickly and easily generate appropriate scripts and commands. Even without technical knowledge, complex commands can be generated simply by inputting instructions in natural language. Therefore, users can perform their tasks efficiently, and because they can visually confirm the generated scripts and commands, they can proceed with confidence.

[0137] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0138] Step 1:

[0139] The user inputs instructions into the terminal using natural language. Specifically, the user inputs text data such as, "Please generate a command to connect to the specified server and check disk usage." This input becomes the input data required to proceed to the next step.

[0140] Step 2:

[0141] The terminal sends the user's input to the server. The terminal sends the entered text as a string to the server, which is then processed by the parsing engine. In this step, the input text is transferred from the terminal to the server.

[0142] Step 3:

[0143] The server analyzes the user's input. The server uses a natural language processing engine to analyze the input and understand the user's intent. In this process, the input data (natural language text) is syntactically analyzed and semantically. The analysis engine extracts the intent "check disk usage" from the input text and generates analysis results to proceed to the next step.

[0144] Step 4:

[0145] The server generates a script or command based on the analyzed data. Here, the server uses a command generation engine to retrieve the appropriate template from the template database and generate a specific command. For example, if the analysis result is "check disk usage", the command "df -h" will be generated. Based on the analysis result (parsed data), the generation engine processes the data and generates a specific command as output.

[0146] Step 5:

[0147] The server sends the generated scripts and commands back to the terminal. Specifically, the server sends the command "df -h" generated by the generation engine to the terminal, and a process of receiving it takes place.

[0148] Step 6:

[0149] The terminal presents the generated script or command to the user. The terminal displays the received command in the user interface, allowing the user to review and use it. For example, the command "df -h" might appear on the screen, and the user can copy and execute that command.

[0150] (Application Example 1)

[0151] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0152] Traditionally, supporting the operation of factory robots often requires advanced technical knowledge and specialized skills to properly execute complex operations and maintenance instructions. This presents challenges, particularly in responding quickly and accurately to emergencies and non-routine tasks. It also increases the risk of operational errors and work delays.

[0153] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0154] In this invention, the server includes an input means for the user to input information in natural language, an analysis means for analyzing the user's input, a generation means for generating a script or command based on the analyzed content, a presentation means for presenting the generated script or command to the user, and an execution means for applying the generated script or command to a machine. This enables the user to quickly and accurately operate and maintain factory robots without requiring specialized knowledge.

[0155] "Input means" refers to functions and devices that allow users to input commands and instructions in natural language.

[0156] "Analysis means" refers to functions and devices for analyzing the content of natural language obtained from input means and understanding the user's intent.

[0157] "Generation means" refers to a function and device for generating an appropriate script or command based on the content analyzed by the analysis means.

[0158] "Presentation means" refers to functions and devices for presenting scripts or commands generated by the generation means to the user visually or by other means.

[0159] "Execution means" refers to functions and devices for actually applying the script or command presented by the presentation means to a machine and performing operations or tasks.

[0160] A "generative AI model" is an artificial intelligence model that analyzes natural language and generates appropriate scripts or commands based on the analysis results.

[0161] This invention provides a system that streamlines the operation of factory robots, enabling users to generate and execute appropriate operation scripts and commands simply by inputting instructions in natural language. This system includes the following main components:

[0162] An "input method" refers to a function or device that allows a user to input commands or instructions in natural language. Specifically, this includes devices such as smartphones, tablets, and factory terminals. The user inputs instructions to the device in natural language, such as "Return the robot arm to the home position."

[0163] "Analysis means" refers to functions and devices for analyzing the content of natural language obtained from input means and understanding the user's intent. This system uses a generative AI model (e.g., OpenAI® GPT-3®) as the analysis engine. This model analyzes the input natural language using advanced algorithms and extracts the intended operation.

[0164] The "generation means" refers to a function and device for generating an appropriate script or command based on the content analyzed by the analysis means. If the analyzed content is an instruction to "return the robot arm to the home position," the generation AI model will generate a script such as "moveToHomePosition();". The generation means is mainly executed on a server, and the generated script is processed within the server.

[0165] "Presentation means" refers to functions and devices for presenting the script or command generated by the generation means to the user visually or by other means. Specifically, it displays the generated script on the screen of a smartphone or tablet. The user can then review its contents and decide whether or not to execute it.

[0166] "Execution means" refers to the functions and devices that actually apply the script or command presented by the presentation means to the machine and perform the operation or task. Once the user confirms the displayed script and approves its execution, the script is sent to the factory robot and the actual operation begins.

[0167] Specific example

[0168] For example, if a user enters "Get the current position of the robot arm" into a tablet, the system's analysis mechanism analyzes the input and uses a generation AI model to generate the command "getCurrentPosition();". The generated command is then displayed on the tablet screen for the user to confirm. After the user confirms, the command is sent to the robot, which obtains its current position and returns that information.

[0169] Example of a prompt

[0170] The following are examples of specific prompt statements used for generative AI models.

[0171] Based on the following natural language instructions, generate a script to operate a factory robot:

[0172] Instructions: Return the robot arm to its home position.

[0173] script:

[0174] In this way, by allowing users to input instructions in natural language and automatically generating appropriate scripts and commands based on those instructions, the operation of factory robots is simplified, and efficient work is achieved.

[0175] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0176] Step 1:

[0177] The user inputs a natural language instruction, such as "Return the robot arm to the home position," into an input device such as a tablet or smartphone. This input information is then transmitted to the system's input method.

[0178] Input: User's natural language instructions

[0179] Output: Instruction data to the input means

[0180] Step 2:

[0181] Natural language instructions received through the input device are sent to the server. The server's analysis device inputs these instructions into a generative AI model (e.g., OpenAI GPT-3) and analyzes the content. The generative AI model analyzes the input natural language and processes the data to understand the user's intent.

[0182] Input: Natural language instruction data

[0183] Output: Analyzed instruction data

[0184] Step 3:

[0185] The generating AI model generates an appropriate script or command based on the analyzed instruction data. For example, in response to the instruction "Return the robot arm to the home position," the script "moveToHomePosition();" is generated. This generated script is then processed by the generation mechanism.

[0186] Input: Analyzed instruction data

[0187] Output: Generated script or command

[0188] Step 4:

[0189] The generated script or command is visually displayed to the user through a presentation mechanism. The user reviews the script on their tablet or smartphone screen. At this point, the user reviews the content of the generated command and decides whether or not to execute it.

[0190] Input: Generated script or command

[0191] Output: Displayed script

[0192] Step 5:

[0193] After the user reviews the presented script and presses the execute button, the script is sent to the factory robot via the server. This causes the robot to begin its actual operation via the execution mechanism. For example, the robot arm might perform an action to return to its home position.

[0194] Input: User verification and execution command

[0195] Output: Actual operation by the robot

[0196] In this process, users can quickly and accurately operate factory robots simply by inputting instructions in natural language. By coordinating the generative AI model, presentation method, and execution method, it is possible to significantly reduce the burden on the user and achieve efficient work.

[0197] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0198] Overall Overview

[0199] The present invention provides a system that automatically generates and presents necessary scripts and commands simply by having the user input instructions in natural language. Furthermore, the present invention incorporates an emotion engine that recognizes the user's emotions, enabling it to provide appropriate responses according to the user's feelings. This system provides support for users to carry out their tasks quickly and accurately.

[0200] System Configuration

[0201] This system includes the following main components:

[0202] 1. Input method: The user inputs using natural language.

[0203] 2. Analysis method: Analyzes the user's input.

[0204] 3. Generation method: Generate a script or command based on the analyzed content.

[0205] 4. Presentation method: The generated script or command is presented to the user.

[0206] 5. Emotion Engine: Recognizes the user's emotions and provides feedback to analysis and generation methods.

[0207] Details of the example

[0208] 1. User Input and Sentiment Recognition

[0209] The user inputs instructions to the terminal using natural language. For example, if the user inputs, "I want to generate a command to connect to a specified server and check disk usage," the emotion engine simultaneously recognizes the user's emotions (e.g., frustration or tension) from the user's input, voice, facial expressions, and gestures.

[0210] 2. Analysis of Natural Language and Sentiment

[0211] The terminal sends user input and emotional information to the server. The server uses a natural language processing engine to analyze the input and an emotional processing engine to analyze the user's emotions. The analysis results include the user's desired objective (e.g., "check disk usage") and the user's emotional state.

[0212] 3. Script / command generation and emotion reflection

[0213] Based on the analyzed data, the server generates appropriate scripts and commands. For example, it generates the Unix command "df -h" to check disk usage. Simultaneously, the emotion engine generates additional information and warning messages to alleviate user frustration and tension (e.g., "Executing this command will show you your disk usage").

[0214] 4. Presentation of results and implementation

[0215] The server generates scripts, commands, and emotionally sensitive messages, which are then sent back to the terminal. The terminal presents these to the user. The user can then review the generated commands and additional information and use them.

[0216] Specific example

[0217] Example 1: Checking disk usage

[0218] User input: The user instructs the terminal to "generate a command to check the server's disk usage."

[0219] Server analysis and generation: The server analyzes the user's instructions and generates the disk usage command "df -h". The emotion engine detects the user's tension and generates an additional message: "Executing this command will show you your disk usage."

[0220] Prompt: The terminal prompts the user with the command "df -h" and an additional message. The user uses the provided command to check the server's disk usage.

[0221] Example 2: Retrieving a file list

[0222] User input: The user instructs the terminal to "generate a command to retrieve a list of files in the specified directory."

[0223] Server analysis and generation: The server analyzes the data, such as "ls -la / path / to / directory", and generates the appropriate command. The sentiment engine detects user frustration and generates an additional message such as, "Executing this command will show you a list of files in the specified directory."

[0224] Presentation: The terminal presents the user with the additional message "ls -la / path / to / directory", and the user executes the command to obtain a list of files.

[0225] Operation flow and effects

[0226] This system allows users to quickly and easily generate appropriate scripts and commands, and receive responses that take their emotions into consideration thanks to its emotion engine. Users not only input instructions in natural language, but the system also enhances the user experience by providing additional information and warning messages based on the user's emotions. Even those lacking technical knowledge can use the system without difficulty, allowing for confident operation.

[0227] The following describes the processing flow.

[0228] Step 1:

[0229] The user enters instructions into the terminal using natural language. For example, they might enter, "I want to generate a command to connect to a specified server and check disk usage."

[0230] Step 2:

[0231] The emotion engine recognizes the user's emotions from their input, voice, facial expressions, and other factors. For example, it can detect user frustration based on the tone of their voice and input speed.

[0232] Step 3:

[0233] The device sends user instructions in natural language and emotional information to the server. The data sent to the server consists of natural language text data and emotional information generated by an emotion engine.

[0234] Step 4:

[0235] The server uses a natural language processing engine to analyze the user's instructions. The processing engine identifies the user's desired objective (for example, "check disk usage").

[0236] Step 5:

[0237] The emotion engine analyzes the transmitted emotion information to identify the user's emotional state. For example, it determines whether the user is feeling "stressed" or "frustrated."

[0238] Step 6:

[0239] The server generates an appropriate script or command based on the analyzed information. For example, it might generate the Unix command "df -h" to check disk usage.

[0240] Step 7:

[0241] The server generates additional messages tailored to the user's emotions based on the analysis results of the emotion engine. For example, if the user is feeling frustrated, it will generate encouraging and reassuring messages such as, "Running this command will show you your disk usage."

[0242] Step 8:

[0243] The server sends the generated script or command and any additional messages to the terminal.

[0244] Step 9:

[0245] The terminal displays the received script or command and any additional messages to the user. The user reviews the displayed content on the terminal screen.

[0246] Step 10:

[0247] If a user needs to execute a presented script or command, they send the command to the server via the terminal. The server executes the received command and returns the result to the user.

[0248] Step 11:

[0249] The terminal displays the execution results of commands received from the server to the user. The user can then obtain the necessary information.

[0250] Through the processing steps described above, users can efficiently perform a series of operations, from natural language input to the generation of appropriate scripts and commands, to emotionally sensitive responses and execution. This system allows users to carry out their work quickly and easily, and the emotional engine provides a less stressful user experience.

[0251] (Example 2)

[0252] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0253] Current scripting and command generation systems require users to possess a high level of technical knowledge and have limited understanding of natural language instructions, thus limiting user convenience. Furthermore, they fail to provide responses that consider the user's emotional state, resulting in a lack of appropriate support, especially for users experiencing frustration or anxiety. Therefore, there is a need for a system that can quickly and accurately understand user instructions and respond with consideration for the user's emotions.

[0254] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0255] In this invention, the server includes an input means for the user to input in natural language, an analysis means for analyzing the user's input and emotional information, a generation means for generating a script or command based on the analyzed content and emotional information, and a presentation means for presenting the generated script or command and additional emotionally sensitive information to the user. As a result, the user can quickly and accurately generate the necessary scripts and commands simply by inputting instructions in natural language, and the system provides responses that are sensitive to the user's emotions, thereby improving the user experience.

[0256] "Input means" refers to a device or interface for a user to input instructions in natural language.

[0257] "Analysis means" refers to a device or software for analyzing user input and emotional information.

[0258] "Generation means" refers to a device or software for generating scripts or commands based on analyzed content and emotional information.

[0259] "Presentation means" refers to a device or interface for presenting a generated script or command and additional information that takes emotions into consideration to the user.

[0260] A "natural language processing engine" is software or an algorithm that analyzes natural language input by a user to understand its intent and purpose.

[0261] An "emotion analysis engine" is software or an algorithm that analyzes a user's emotional state from inputs such as voice, facial expressions, and gestures.

[0262] A "script" is a set of commands or program code generated to automate a specific task or operation.

[0263] A "command" is an instruction generated to tell a system to perform a specific operation or task.

[0264] "Additional information" refers to supplementary messages or explanations provided alongside generated scripts or commands, which are designed with the user's feelings in mind.

[0265] Overall Overview

[0266] The system of this invention automatically generates and presents necessary scripts and commands simply by the user inputting in natural language. Furthermore, the invention incorporates an emotion engine that recognizes the user's emotions, enabling it to provide appropriate responses according to the user's feelings. This system provides support that allows users to carry out their tasks quickly and accurately, even if they lack technical knowledge.

[0267] System Configuration

[0268] This system includes the following main components:

[0269] 1. Input means: A device or interface in which the user inputs information in natural language.

[0270] 2. Analysis means: A device or software that analyzes user input and emotional information.

[0271] 3. Generation means: A device or software that generates a script or command based on the analyzed content and emotional information.

[0272] 4. Presentation means: A device or interface that presents the generated script or command and additional information that takes emotions into consideration to the user.

[0273] Hardware and software to be used

[0274] The following hardware and software are used for the specific implementation of the system:

[0275] Natural language analysis engine: For example, use software such as NLTK or OpenAI's GPT-4.

[0276] Sentiment analysis engine: For example, use software such as Google's (registered trademark) Dialogflow.

[0277] Terminal: Computers, smartphones, etc. directly operated by the user.

[0278] Server: A remote server for executing analysis means and generation means.

[0279] Specific operation examples of the system

[0280] User input and emotion recognition

[0281] The user inputs instructions in natural language to the terminal. For example, input "I want to generate a command to connect to a specified server and check the disk usage status". Also, the terminal recognizes the user's emotion (e.g., frustration or tension) from the user's input content, voice, expression, and gesture.

[0282] Analysis of natural language and emotion

[0283] The terminal sends the user's input content and emotion information to the server. The server analyzes the input content using a natural language analysis engine and analyzes the user's emotion using an emotion engine. The analysis results include the purpose the user wants to achieve (e.g., "checking the disk usage status") and the user's emotional state.

[0284] Script / command generation and emotion reflection

[0285] Based on the parsed content, the server generates appropriate scripts and commands. For example, it generates the Unix command "df -h" to check disk usage. At the same time, the emotion engine generates additional information (such as "Executing this command will show you the disk usage") to relieve the user's frustration and tension.

[0286] Result Presentation and Execution

[0287] The scripts, commands generated by the server, and the emotion - considerate messages are sent back to the terminal. The terminal presents these to the user. The user can check the generated commands and additional information and use them.

[0288] Examples of Prompt Sentences

[0289] Examples of specific prompt sentences are shown below:

[0290] "Please generate a Unix command to check the disk usage by connecting to the specified server. Considering that the user is nervous, also add an explanation of the execution result."

[0291] "Please generate a Unix command to get a list of files in the specified directory. If the user is feeling frustrated, also include a message explaining how this command can be useful."

[0292] In this way, the present invention realizes a system that flexibly and quickly analyzes the user's input in natural language and provides an appropriate response considering emotions.

[0293] The flow of specific processing in Example 2 will be described using FIG. 13.

[0294] Step 1:

[0295] The user makes an input in natural language

[0296] The user inputs an instruction in natural language towards the terminal. For example, the user inputs "Generate a command to check the disk usage status of the server". This input content is saved as text data in the terminal.

[0297] Step 2:

[0298] The terminal acquires the input content and emotion information.

[0299] In addition to the user's input content, the terminal also acquires emotion information such as the user's voice, expression, and gesture. As a result, the input data includes the user's text data and emotion data.

[0300] Step 3:

[0301] The terminal sends the input data to the server.

[0302] The terminal sends the collected text data and emotion data to the server. This data is packaged in a format such as JSON and transferred to the server via the network.

[0303] Step 4:

[0304] The server analyzes the natural language.

[0305] The server analyzes the received text data using a natural language analysis engine (such as NLTK or OpenAI's GPT-4). Specifically, it analyzes the user's instruction content to understand the purpose of "checking the disk usage status". This analysis result is generated as intermediate data.

[0306] Step 5:

[0307] The server analyzes the emotion information.

[0308] The server uses an emotion analysis engine (Google's Dialogflow) to analyze emotional data. For example, it can identify if a user is feeling nervous based on their voice tone and facial expressions. This analysis result is also generated as intermediate data.

[0309] Step 6:

[0310] The server generates commands or scripts.

[0311] The server integrates natural language processing results and sentiment analysis results to generate appropriate scripts and commands. For example, it generates the Unix command "df -h" to check disk usage. This generated command is then output.

[0312] Step 7:

[0313] The server generates additional information that takes emotions into consideration.

[0314] Based on the sentiment analysis results, the server generates additional information and explanations that take the user's emotions into consideration. For example, it might generate a message such as, "Executing this command will show you your disk usage." The command message with this additional information is then output.

[0315] Step 8:

[0316] The server sends the generated data to the terminal.

[0317] The server sends the generated scripts, commands, and additional information to the terminal. This data is also structured in JSON format or similar and transferred to the terminal over the network.

[0318] Step 9:

[0319] The device presents to the user

[0320] The terminal presents the received scripts, commands, and any additional messages to the user. The user reviews them and, if necessary, executes the commands directly on the terminal.

[0321] (Application Example 2)

[0322] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0323] Conventional factory robot systems require staff to perform complex coding and command input when issuing operating instructions to the robots, thus demanding specialized knowledge. Furthermore, a lack of consideration for staff emotional states often leads to stress and frustration. There is a need to solve these problems and make factory robot operation more efficient and user-friendly.

[0324] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an input means for the user to input in natural language, an analysis means for analyzing the user's input content and emotional state, a generation means for generating a script or command based on the analyzed content and emotional state, and a presentation means for presenting the generated script or command and an emotionally sensitive response message to the user. As a result, factory staff can easily instruct robot operation in natural language, and emotionally sensitive responses are provided, thereby reducing staff stress and enabling efficient and smooth operation.

[0325] "Input means" refers to a device or interface for a user to input instructions in natural language.

[0326] "Analysis means" refers to a device or program for analyzing the user's input and emotional state.

[0327] "Generation means" refers to a device or program for generating a script or command based on the analyzed content and emotional state.

[0328] "Presentation means" refers to a device or interface for presenting a generated script or command and an emotionally sensitive response message to the user.

[0329] An "emotion recognition engine" is a program or device that analyzes a user's facial expressions and voice tone to identify their emotional state.

[0330] A "natural language processing engine" is a program or device used to analyze the content of natural language input by a user.

[0331] A "server" is a central processing unit that executes analysis and generation methods and provides the results to the user.

[0332] A "script" is a series of instructions generated to automate a specific task.

[0333] A "command" is an instruction issued to perform a specific operation or task.

[0334] A "response message" is additional information or a warning message provided with consideration for the user's emotional state.

[0335] This invention is a system in which a user inputs instructions in natural language, generates a script or command based on the analyzed content, and then presents the generated result. The invention incorporates an emotion recognition engine that recognizes the user's emotions, thereby enabling it to provide appropriate responses according to the user's emotions.

[0336] System Overview

[0337] This system includes the following main components:

[0338] 1. Input method: The user inputs information using natural language. This includes smartphone apps and factory robot interfaces.

[0339] 2. Analysis Method: The user's input and emotional state are analyzed. This includes the Google Cloud Natural Language API and the emotion recognition engine.

[0340] 3. Generation method: Generates a script or command based on the analyzed content and emotional state. OpenAI's generative AI model is an example of this.

[0341] 4. Presentation method: The generated script or command and an emotionally sensitive response message are presented to the user. This includes smartphone apps and robot display devices.

[0342] Program processing

[0343] The server first receives natural language instructions entered by the user. Next, it uses an emotion recognition engine to analyze the user's facial expressions and tone of voice to identify their emotional state. The analysis method uses the Google Cloud Natural Language API to analyze and understand the user's instructions.

[0344] In the generation process, OpenAI's generation AI model generates appropriate scripts or commands based on the analysis results and emotional state. Simultaneously, response messages corresponding to the user's emotional state are also generated. For example, if the user is feeling stressed, a reassuring message will be generated.

[0345] The generated script or command and response messages are presented to the user via a smartphone app or the robot's display device. The user can then proceed with their task using the presented script or command.

[0346] Specific examples and prompt statements

[0347] For example, consider a scenario where a user instructs a factory line to stop.

[0348] User instruction: "Stop factory line 1."

[0349] Example of a prompt:

[0350] Effect: Stop factory line 1

[0351] Emotion Score: -0.5

[0352] Generate the appropriate command and supportive message.

[0353] In this case, the server generates a response message along with the command "stop line 1" that reads, "The operation is complete. Please let us know if there is anything else we can help you with."

[0354] This system allows users to intuitively control robots using natural language without having to input complex commands, enabling them to work efficiently and without experiencing stress during the process.

[0355] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0356] Step 1:

[0357] The user inputs instructions. The user inputs instructions in natural language using a smartphone app or the factory robot's interface. The entered natural language data is sent to the server.

[0358] Step 2:

[0359] The server receives the input. The server receives natural language data sent by the user and stores its contents. The input is natural language text data. The output is the storage and management of this text data.

[0360] Step 3:

[0361] The system performs emotion analysis using an emotion recognition engine. The server passes the input natural language data to the Google Cloud Natural Language API, which analyzes the user's emotional state from voice tone and facial expression data (which may be obtained via webcam or microphone). The input consists of natural language text data and voice / facial expression data, and the output is an emotional state score.

[0362] Step 4:

[0363] This system performs natural language processing. The server passes input data to the Google Cloud Natural Language API for language analysis. The input is natural language text data, and the output is the analyzed content (text structure and semantic information as a result of the analysis).

[0364] Step 5:

[0365] The server generates scripts and response messages. It passes the analysis results and emotional state scores to OpenAI's generative AI model, which then generates appropriate scripts or commands and emotionally sensitive response messages. The input is the analysis results and emotional state scores, and the output is the generated scripts or commands and response messages.

[0366] Step 6:

[0367] The generated results are presented. The server sends the generated script or command and response message to the user's terminal. The user receives the presented results through a smartphone app or a robot's display device. The input is the generated script or command and response message, and the output is the display on the user interface.

[0368] Step 7:

[0369] User verification and execution. The user reviews the presented script or command and response messages and uses them to proceed with the task. The user provides further input as needed, and the process continues.

[0370] These steps enable users to easily generate and execute scripts or commands based on emotionally sensitive natural language instructions.

[0371] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0372] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0373] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0374] [Second Embodiment]

[0375] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0376] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0377] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0378] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0379] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0380] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0381] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0382] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0383] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0384] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0385] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0386] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0387] Overall Overview

[0388] The present invention provides a system that automatically generates and presents necessary scripts and commands simply by having the user input instructions in natural language. This system supports users in carrying out their tasks quickly and accurately.

[0389] System Configuration

[0390] This system includes the following main components:

[0391] 1. Input method: The user inputs using natural language.

[0392] 2. Analysis method: Analyzes the user's input.

[0393] 3. Generation method: Generate a script or command based on the analyzed content.

[0394] 4. Presentation method: The generated script or command is presented to the user.

[0395] Details of the example

[0396] 1. User input

[0397] The user inputs instructions into the terminal using natural language. For example, they might input, "I want to generate a command to connect to a specified server and check disk usage."

[0398] 2. Natural language analysis

[0399] The terminal sends user input to the server. The server analyzes the input using a natural language processing engine. The processing engine understands what the user wants to achieve and extracts information to generate appropriate scripts or commands.

[0400] 3. Script / command generation

[0401] Based on the analyzed information, the server generates appropriate scripts and commands. For example, the server generates the Unix command "df -h" to check disk usage.

[0402] 4. Presentation of results

[0403] The server generates scripts and commands, which are then sent back to the terminal. The terminal then presents these to the user. The user can review the generated commands and use them.

[0404] Specific example

[0405] Example 1: Checking disk usage

[0406] User input: The user instructs the terminal to "generate a command to check the server's disk usage."

[0407] Server analysis and generation: The server analyzes the user's instructions and generates the disk usage command "df -h".

[0408] Prompt: The terminal prompts the user to execute "df -h". The user uses the suggested command to check the server's disk usage.

[0409] Example 2: Retrieving a file list

[0410] User input: The user instructs the terminal to "generate a command to retrieve a list of files in the specified directory."

[0411] Server analysis and generation: The server analyzes the data using commands like "ls -la / path / to / directory" and generates the appropriate commands.

[0412] Presentation: The terminal presents the user with the command "ls -la / path / to / directory", and the user executes the command to obtain a list of files.

[0413] Operation flow and effects

[0414] This system allows users to quickly and easily generate appropriate scripts and commands, enabling them to work efficiently. Because users can generate complex commands simply by inputting instructions in natural language, the system can be used without problems even by those lacking technical knowledge. Furthermore, the generated scripts and commands can be visually reviewed, allowing users to proceed with confidence.

[0415] The following describes the processing flow.

[0416] Step 1:

[0417] The user enters instructions into the terminal using natural language. For example, they might enter, "I want to generate a command to connect to a specified server and check disk usage."

[0418] Step 2:

[0419] The terminal receives user input and sends that input to the server. The data sent to the server is natural language text data.

[0420] Step 3:

[0421] The server sends the data to a natural language processing engine to analyze the user's input. The natural language processing engine analyzes the input sentence and extracts information to understand the user's intent.

[0422] Step 4:

[0423] The natural language processing engine sends the analysis results back to the server. The analysis results include the user's desired objective (for example, "check disk usage").

[0424] Step 5:

[0425] The server generates an appropriate script or command based on the analysis results. For example, it generates the command "df -h" to check disk usage.

[0426] Step 6:

[0427] The server sends the generated script or command back to the terminal. The terminal receives this data.

[0428] Step 7:

[0429] The terminal displays the received script or command to the user. The user confirms the displayed content on the terminal screen.

[0430] Step 8:

[0431] If a user needs to execute a presented script or command, they send the command to the server via the terminal. The server executes the received command and returns the result to the user.

[0432] Step 9:

[0433] The terminal displays the execution results of commands received from the server to the user. The user can then obtain the necessary information.

[0434] Through the above processing steps, users can efficiently perform a series of operations, from natural language input to the generation and execution of appropriate scripts and commands. This system makes it easy to generate and execute complex commands, even for those lacking technical knowledge.

[0435] (Example 1)

[0436] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0437] In conventional systems, users needed knowledge of commands and scripts to perform the desired operations, making them difficult for non-specialized users. Furthermore, generating scripts and commands had to be done manually, which was time-consuming and laborious. Therefore, there was a need for a system that would allow users to easily generate commands using natural language.

[0438] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0439] In this invention, the server includes an input means for the user to input in natural language, a transmission means for the terminal to send the user's input to the server, an analysis means for the server to analyze the user's input, a generation means to generate a script or command based on the analyzed content, and a presentation means to return the script or command generated from the server to the terminal and present it to the user. This makes it possible for the user to easily generate the desired script or command in natural language and to execute operations quickly and accurately.

[0440] "Natural language" refers to the language that humans use on a daily basis, and is not a specific programming language or specialized terminology, but rather a general language.

[0441] "Input method" refers to a device or interface that allows a user to input information or instructions in natural language.

[0442] "Transmission means" refers to the communication functions and protocols that a terminal uses to send user input to a server.

[0443] "Analysis means" refers to algorithms and engines that semantically understand the natural language input received by the server and perform appropriate processing.

[0444] A "natural language processing engine" refers to software or a system that analyzes input natural language and understands its meaning and intent.

[0445] "Generation means" refers to algorithms or engines used to generate specific scripts or commands based on the analyzed content.

[0446] A "template database" refers to a set of predefined templates that a generation tool references when generating scripts or commands.

[0447] "Presentation means" refers to devices or interfaces used to display generated scripts or commands to the user.

[0448] A "terminal" refers to a computer or mobile device used by a user to input information and receive generated scripts or commands.

[0449] A "server" refers to a computer system equipped with analysis and generation capabilities, designed to process input from terminals.

[0450] Overall Overview

[0451] The present invention provides a system that automatically generates and presents necessary scripts and commands simply by having the user input instructions in natural language. This allows users to perform operations efficiently even without specialized knowledge. The system provides support for users to carry out their tasks quickly and accurately.

[0452] System Configuration

[0453] This system includes the following main components:

[0454] 1. Input method: The user inputs using natural language.

[0455] 2. Transmission method: The terminal sends the user's input to the server.

[0456] 3. Analysis method: The server analyzes the user's input.

[0457] 4. Generation method: Generate a script or command based on the analyzed content.

[0458] 5. Presentation method: The generated script or command is presented to the user.

[0459] Details of the example

[0460] 1. User input

[0461] The user inputs instructions into the terminal using natural language. For example, they might input, "Please generate a command to connect to the specified server and check disk usage." An example of a specific prompt is as follows:

[0462] Example of a prompt:

[0463] "Please generate a command to connect to the specified server and check disk usage."

[0464] "Generate a command to retrieve a list of files in a specified directory."

[0465] 2. Natural language analysis

[0466] The terminal sends the user's input to the server. The server analyzes the input using a natural language processing engine (e.g., "GPT-4"). The natural language processing engine syntactically analyzes the input text, understands its intent, and then extracts the necessary information.

[0467] 3. Script / command generation

[0468] Based on the analyzed data, the server generates appropriate scripts and commands. In this process, the server uses a command generation engine (e.g., a "shell script generation library") to generate specific commands by supplementing templates obtained from a template database. For example, to check disk usage, the Unix command "df -h" is generated.

[0469] 4. Presentation of results

[0470] The generated scripts and commands are sent back from the server to the terminal. The terminal visually presents them to the user, who can then review and use them. For example, the generated command "df -h" is displayed on the terminal screen and executed by the user.

[0471] Specific example

[0472] Example 1: Checking disk usage

[0473] User input: "Generate a command to check the server's disk usage."

[0474] Server analysis and generation: The server analyzes the input and generates the command "df -h".

[0475] Instructions: The terminal will display "df -h" to the user, who will use this to check disk usage.

[0476] Example 2: Retrieving a file list

[0477] User input: "Generate a command to get a list of files in the specified directory."

[0478] Server analysis and generation: The server generates the command "ls -la / path / to / directory".

[0479] Instructions: The terminal displays "ls -la / path / to / directory" to the user, who then uses this to obtain a list of files.

[0480] Operation flow and effects

[0481] This system allows users to quickly and easily generate appropriate scripts and commands. Even without technical knowledge, complex commands can be generated simply by inputting instructions in natural language. Therefore, users can perform their tasks efficiently, and because they can visually confirm the generated scripts and commands, they can proceed with confidence.

[0482] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0483] Step 1:

[0484] The user inputs instructions into the terminal using natural language. Specifically, the user inputs text data such as, "Please generate a command to connect to the specified server and check disk usage." This input becomes the input data required to proceed to the next step.

[0485] Step 2:

[0486] The terminal sends the user's input to the server. The terminal sends the entered text as a string to the server, which is then processed by the parsing engine. In this step, the input text is transferred from the terminal to the server.

[0487] Step 3:

[0488] The server analyzes the user's input. The server uses a natural language processing engine to analyze the input and understand the user's intent. In this process, the input data (natural language text) is syntactically analyzed and semantically. The analysis engine extracts the intent "check disk usage" from the input text and generates analysis results to proceed to the next step.

[0489] Step 4:

[0490] The server generates a script or command based on the analyzed data. Here, the server uses a command generation engine to retrieve the appropriate template from the template database and generate a specific command. For example, if the analysis result is "check disk usage", the command "df -h" will be generated. Based on the analysis result (parsed data), the generation engine processes the data and generates a specific command as output.

[0491] Step 5:

[0492] The server sends the generated scripts and commands back to the terminal. Specifically, the server sends the command "df -h" generated by the generation engine to the terminal, and a process of receiving it takes place.

[0493] Step 6:

[0494] The terminal presents the generated script or command to the user. The terminal displays the received command in the user interface, allowing the user to review and use it. For example, the command "df -h" might appear on the screen, and the user can copy and execute that command.

[0495] (Application Example 1)

[0496] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0497] Traditionally, supporting the operation of factory robots often requires advanced technical knowledge and specialized skills to properly execute complex operations and maintenance instructions. This presents challenges, particularly in responding quickly and accurately to emergencies and non-routine tasks. It also increases the risk of operational errors and work delays.

[0498] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0499] In this invention, the server includes an input means for the user to input information in natural language, an analysis means for analyzing the user's input, a generation means for generating a script or command based on the analyzed content, a presentation means for presenting the generated script or command to the user, and an execution means for applying the generated script or command to a machine. This enables the user to quickly and accurately operate and maintain factory robots without requiring specialized knowledge.

[0500] "Input means" refers to functions and devices that allow users to input commands and instructions in natural language.

[0501] "Analysis means" refers to functions and devices for analyzing the content of natural language obtained from input means and understanding the user's intent.

[0502] "Generation means" refers to a function and device for generating an appropriate script or command based on the content analyzed by the analysis means.

[0503] "Presentation means" refers to functions and devices for presenting scripts or commands generated by the generation means to the user visually or by other means.

[0504] "Execution means" refers to functions and devices for actually applying the script or command presented by the presentation means to a machine and performing operations or tasks.

[0505] A "generative AI model" is an artificial intelligence model that analyzes natural language and generates appropriate scripts or commands based on the analysis results.

[0506] This invention provides a system that streamlines the operation of factory robots, enabling users to generate and execute appropriate operation scripts and commands simply by inputting instructions in natural language. This system includes the following main components:

[0507] An "input method" refers to a function or device that allows a user to input commands or instructions in natural language. Specifically, this includes devices such as smartphones, tablets, and factory terminals. The user inputs instructions to the device in natural language, such as "Return the robot arm to the home position."

[0508] "Analysis means" refers to functions and devices for analyzing the content of natural language obtained from input means and understanding the user's intent. This system uses a generative AI model (e.g., OpenAI GPT-3) as the analysis engine. This model analyzes the input natural language using advanced algorithms and extracts the intended operation.

[0509] The "generation means" refers to a function and device for generating an appropriate script or command based on the content analyzed by the analysis means. If the analyzed content is an instruction to "return the robot arm to the home position," the generation AI model will generate a script such as "moveToHomePosition();". The generation means is mainly executed on a server, and the generated script is processed within the server.

[0510] "Presentation means" refers to functions and devices for presenting the script or command generated by the generation means to the user visually or by other means. Specifically, it displays the generated script on the screen of a smartphone or tablet. The user can then review its contents and decide whether or not to execute it.

[0511] "Execution means" refers to the functions and devices that actually apply the script or command presented by the presentation means to the machine and perform the operation or task. Once the user confirms the displayed script and approves its execution, the script is sent to the factory robot and the actual operation begins.

[0512] Specific example

[0513] For example, if a user enters "Get the current position of the robot arm" into a tablet, the system's analysis mechanism analyzes the input and uses a generation AI model to generate the command "getCurrentPosition();". The generated command is then displayed on the tablet screen for the user to confirm. After the user confirms, the command is sent to the robot, which obtains its current position and returns that information.

[0514] Example of a prompt

[0515] The following are examples of specific prompt statements used for generative AI models.

[0516] Based on the following natural language instructions, generate a script to operate a factory robot:

[0517] Instructions: Return the robot arm to its home position.

[0518] script:

[0519] In this way, by allowing users to input instructions in natural language and automatically generating appropriate scripts and commands based on those instructions, the operation of factory robots is simplified, and efficient work is achieved.

[0520] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0521] Step 1:

[0522] The user inputs a natural language instruction, such as "Return the robot arm to the home position," into an input device such as a tablet or smartphone. This input information is then transmitted to the system's input method.

[0523] Input: User's natural language instructions

[0524] Output: Instruction data to the input means

[0525] Step 2:

[0526] Natural language instructions received through the input device are sent to the server. The server's analysis device inputs these instructions into a generative AI model (e.g., OpenAI GPT-3) and analyzes the content. The generative AI model analyzes the input natural language and processes the data to understand the user's intent.

[0527] Input: Natural language instruction data

[0528] Output: Analyzed instruction data

[0529] Step 3:

[0530] The generating AI model generates an appropriate script or command based on the analyzed instruction data. For example, in response to the instruction "Return the robot arm to the home position," the script "moveToHomePosition();" is generated. This generated script is then processed by the generation mechanism.

[0531] Input: Analyzed instruction data

[0532] Output: Generated script or command

[0533] Step 4:

[0534] The generated script or command is visually displayed to the user through a presentation mechanism. The user reviews the script on their tablet or smartphone screen. At this point, the user reviews the content of the generated command and decides whether or not to execute it.

[0535] Input: Generated script or command

[0536] Output: Displayed script

[0537] Step 5:

[0538] After the user reviews the presented script and presses the execute button, the script is sent to the factory robot via the server. This causes the robot to begin its actual operation via the execution mechanism. For example, the robot arm might perform an action to return to its home position.

[0539] Input: User verification and execution command

[0540] Output: Actual operation by the robot

[0541] In this process, users can quickly and accurately operate factory robots simply by inputting instructions in natural language. By coordinating the generative AI model, presentation method, and execution method, it is possible to significantly reduce the burden on the user and achieve efficient work.

[0542] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0543] Overall Overview

[0544] The present invention provides a system that automatically generates and presents necessary scripts and commands simply by having the user input instructions in natural language. Furthermore, the present invention incorporates an emotion engine that recognizes the user's emotions, enabling it to provide appropriate responses according to the user's feelings. This system provides support for users to carry out their tasks quickly and accurately.

[0545] System Configuration

[0546] This system includes the following main components:

[0547] 1. Input method: The user inputs using natural language.

[0548] 2. Analysis method: Analyzes the user's input.

[0549] 3. Generation method: Generate a script or command based on the analyzed content.

[0550] 4. Presentation method: The generated script or command is presented to the user.

[0551] 5. Emotion Engine: Recognizes the user's emotions and provides feedback to analysis and generation methods.

[0552] Details of the example

[0553] 1. User Input and Sentiment Recognition

[0554] The user inputs instructions to the terminal using natural language. For example, if the user inputs, "I want to generate a command to connect to a specified server and check disk usage," the emotion engine simultaneously recognizes the user's emotions (e.g., frustration or tension) from the user's input, voice, facial expressions, and gestures.

[0555] 2. Analysis of Natural Language and Sentiment

[0556] The terminal sends user input and emotional information to the server. The server uses a natural language processing engine to analyze the input and an emotional processing engine to analyze the user's emotions. The analysis results include the user's desired objective (e.g., "check disk usage") and the user's emotional state.

[0557] 3. Script / command generation and emotion reflection

[0558] Based on the analyzed data, the server generates appropriate scripts and commands. For example, it generates the Unix command "df -h" to check disk usage. Simultaneously, the emotion engine generates additional information and warning messages to alleviate user frustration and tension (e.g., "Executing this command will show you your disk usage").

[0559] 4. Presentation of results and implementation

[0560] The server generates scripts, commands, and emotionally sensitive messages, which are then sent back to the terminal. The terminal presents these to the user. The user can then review the generated commands and additional information and use them.

[0561] Specific example

[0562] Example 1: Checking disk usage

[0563] User input: The user instructs the terminal to "generate a command to check the server's disk usage."

[0564] Server analysis and generation: The server analyzes the user's instructions and generates the disk usage command "df -h". The emotion engine detects the user's tension and generates an additional message: "Executing this command will show you your disk usage."

[0565] Prompt: The terminal prompts the user with the command "df -h" and an additional message. The user uses the provided command to check the server's disk usage.

[0566] Example 2: Retrieving a file list

[0567] User input: The user instructs the terminal to "generate a command to retrieve a list of files in the specified directory."

[0568] Server analysis and generation: The server analyzes the data, such as "ls -la / path / to / directory", and generates the appropriate command. The sentiment engine detects user frustration and generates an additional message such as, "Executing this command will show you a list of files in the specified directory."

[0569] Presentation: The terminal presents the user with the additional message "ls -la / path / to / directory", and the user executes the command to obtain a list of files.

[0570] Operation flow and effects

[0571] This system allows users to quickly and easily generate appropriate scripts and commands, and receive responses that take their emotions into consideration thanks to its emotion engine. Users not only input instructions in natural language, but the system also enhances the user experience by providing additional information and warning messages based on the user's emotions. Even those lacking technical knowledge can use the system without difficulty, allowing for confident operation.

[0572] The following describes the processing flow.

[0573] Step 1:

[0574] The user enters instructions into the terminal using natural language. For example, they might enter, "I want to generate a command to connect to a specified server and check disk usage."

[0575] Step 2:

[0576] The emotion engine recognizes the user's emotions from their input, voice, facial expressions, and other factors. For example, it can detect user frustration based on the tone of their voice and input speed.

[0577] Step 3:

[0578] The device sends user instructions in natural language and emotional information to the server. The data sent to the server consists of natural language text data and emotional information generated by an emotion engine.

[0579] Step 4:

[0580] The server uses a natural language processing engine to analyze the user's instructions. The processing engine identifies the user's desired objective (for example, "check disk usage").

[0581] Step 5:

[0582] The emotion engine analyzes the transmitted emotion information to identify the user's emotional state. For example, it determines whether the user is feeling "stressed" or "frustrated."

[0583] Step 6:

[0584] The server generates an appropriate script or command based on the analyzed information. For example, it might generate the Unix command "df -h" to check disk usage.

[0585] Step 7:

[0586] The server generates additional messages tailored to the user's emotions based on the analysis results of the emotion engine. For example, if the user is feeling frustrated, it will generate encouraging and reassuring messages such as, "Running this command will show you your disk usage."

[0587] Step 8:

[0588] The server sends the generated script or command and any additional messages to the terminal.

[0589] Step 9:

[0590] The terminal displays the received script or command and any additional messages to the user. The user reviews the displayed content on the terminal screen.

[0591] Step 10:

[0592] If a user needs to execute a presented script or command, they send the command to the server via the terminal. The server executes the received command and returns the result to the user.

[0593] Step 11:

[0594] The terminal displays the execution results of commands received from the server to the user. The user can then obtain the necessary information.

[0595] Through the processing steps described above, users can efficiently perform a series of operations, from natural language input to the generation of appropriate scripts and commands, to emotionally sensitive responses and execution. This system allows users to carry out their work quickly and easily, and the emotional engine provides a less stressful user experience.

[0596] (Example 2)

[0597] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0598] Current scripting and command generation systems require users to possess a high level of technical knowledge and have limited understanding of natural language instructions, thus limiting user convenience. Furthermore, they fail to provide responses that consider the user's emotional state, resulting in a lack of appropriate support, especially for users experiencing frustration or anxiety. Therefore, there is a need for a system that can quickly and accurately understand user instructions and respond with consideration for the user's emotions.

[0599] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0600] In this invention, the server includes an input means for the user to input in natural language, an analysis means for analyzing the user's input and emotional information, a generation means for generating a script or command based on the analyzed content and emotional information, and a presentation means for presenting the generated script or command and additional emotionally sensitive information to the user. As a result, the user can quickly and accurately generate the necessary scripts and commands simply by inputting instructions in natural language, and the system provides responses that are sensitive to the user's emotions, thereby improving the user experience.

[0601] "Input means" refers to a device or interface for a user to input instructions in natural language.

[0602] "Analysis means" refers to a device or software for analyzing user input and emotional information.

[0603] "Generation means" refers to a device or software for generating scripts or commands based on analyzed content and emotional information.

[0604] "Presentation means" refers to a device or interface for presenting a generated script or command and additional information that takes emotions into consideration to the user.

[0605] A "natural language processing engine" is software or an algorithm that analyzes natural language input by a user to understand its intent and purpose.

[0606] An "emotion analysis engine" is software or an algorithm that analyzes a user's emotional state from inputs such as voice, facial expressions, and gestures.

[0607] A "script" is a set of commands or program code generated to automate a specific task or operation.

[0608] A "command" is an instruction generated to tell a system to perform a specific operation or task.

[0609] "Additional information" refers to supplementary messages or explanations provided alongside generated scripts or commands, which are designed with the user's feelings in mind.

[0610] Overall Overview

[0611] The system of this invention automatically generates and presents necessary scripts and commands simply by the user inputting in natural language. Furthermore, the invention incorporates an emotion engine that recognizes the user's emotions, enabling it to provide appropriate responses according to the user's feelings. This system provides support that allows users to carry out their tasks quickly and accurately, even if they lack technical knowledge.

[0612] System Configuration

[0613] This system includes the following main components:

[0614] 1. Input means: A device or interface in which the user inputs information in natural language.

[0615] 2. Analysis means: A device or software that analyzes user input and emotional information.

[0616] 3. Generation means: A device or software that generates a script or command based on the analyzed content and emotional information.

[0617] 4. Presentation means: A device or interface that presents the generated script or command and additional information that takes emotions into consideration to the user.

[0618] Hardware and software to be used

[0619] The following hardware and software will be used for the specific implementation of the system:

[0620] Natural language processing engine: For example, software such as NLTK or OpenAI's GPT-4 can be used.

[0621] Sentiment analysis engine: For example, software such as Google's Dialogflow can be used.

[0622] Device: A computer, smartphone, or other device that the user directly operates.

[0623] Server: A remote server for executing analysis or generation methods.

[0624] Specific examples of how the system works

[0625] User input and emotion recognition

[0626] The user inputs instructions to the terminal using natural language. For example, they might input, "I want to generate a command to connect to the specified server and check disk usage." The terminal also recognizes the user's emotions (e.g., frustration or tension) from their input, voice, facial expressions, and gestures.

[0627] Analysis of natural language and emotions

[0628] The terminal sends user input and emotional information to the server. The server uses a natural language processing engine to analyze the input and an emotional processing engine to analyze the user's emotions. The analysis results include the user's desired objective (e.g., "check disk usage") and the user's emotional state.

[0629] Script / command generation and emotion reflection

[0630] Based on the analyzed data, the server generates appropriate scripts and commands. For example, it generates the Unix command "df -h" to check disk usage. At the same time, the emotion engine generates additional information to alleviate user frustration and tension (for example, "Executing this command will show you your disk usage").

[0631] Presentation of results and execution

[0632] The server generates scripts, commands, and emotionally sensitive messages, which are then sent back to the terminal. The terminal presents these to the user. The user can then review the generated commands and additional information and use them.

[0633] Example of a prompt

[0634] Here are some examples of specific prompt messages:

[0635] "Please generate Unix commands to connect to the specified server and check disk usage. Considering the user's anxiety, please also include an explanation of the execution results."

[0636] "Generate a Unix command to retrieve a list of files in a specified directory. Include a message explaining how this command can be helpful if the user is experiencing frustration."

[0637] In this way, the present invention realizes a system that flexibly and quickly analyzes the user's natural language input and provides an appropriate response that takes emotions into consideration.

[0638] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0639] Step 1:

[0640] The user enters information in natural language.

[0641] The user inputs instructions into the terminal using natural language. For example, they might input, "Generate a command to check the server's disk usage." This input is saved to the terminal as text data.

[0642] Step 2:

[0643] The device retrieves input content and emotional information.

[0644] The device acquires not only the user's input but also emotional information such as the user's voice, facial expressions, and gestures. As a result, the input data includes both the user's text data and emotional data.

[0645] Step 3:

[0646] The terminal sends the input data to the server.

[0647] The device sends collected text and sentiment data to the server. This data is packaged in JSON format or similar and transferred to the server over the network.

[0648] Step 4:

[0649] The server analyzes natural language.

[0650] The server analyzes the received text data using a natural language processing engine (such as NLTK or OpenAI's GPT-4). Specifically, it analyzes the user's instructions and understands the objective, which is "checking disk usage." This analysis result is then generated as intermediate data.

[0651] Step 5:

[0652] The server analyzes emotional information.

[0653] The server uses an emotion analysis engine (Google's Dialogflow) to analyze emotional data. For example, it can identify if a user is feeling nervous based on their voice tone and facial expressions. This analysis result is also generated as intermediate data.

[0654] Step 6:

[0655] The server generates commands or scripts.

[0656] The server integrates natural language processing results and sentiment analysis results to generate appropriate scripts and commands. For example, it generates the Unix command "df -h" to check disk usage. This generated command is then output.

[0657] Step 7:

[0658] The server generates additional information that takes emotions into consideration.

[0659] Based on the sentiment analysis results, the server generates additional information and explanations that take the user's emotions into consideration. For example, it might generate a message such as, "Executing this command will show you your disk usage." The command message with this additional information is then output.

[0660] Step 8:

[0661] The server sends the generated data to the terminal.

[0662] The server sends the generated scripts, commands, and additional information to the terminal. This data is also structured in JSON format or similar and transferred to the terminal over the network.

[0663] Step 9:

[0664] The device presents to the user

[0665] The terminal presents the received scripts, commands, and any additional messages to the user. The user reviews them and, if necessary, executes the commands directly on the terminal.

[0666] (Application Example 2)

[0667] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0668] Conventional factory robot systems require staff to perform complex coding and command input when issuing operating instructions to the robots, thus demanding specialized knowledge. Furthermore, a lack of consideration for staff emotional states often leads to stress and frustration. There is a need to solve these problems and make factory robot operation more efficient and user-friendly.

[0669] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an input means for the user to input in natural language, an analysis means for analyzing the user's input content and emotional state, a generation means for generating a script or command based on the analyzed content and emotional state, and a presentation means for presenting the generated script or command and an emotionally sensitive response message to the user. As a result, factory staff can easily instruct robot operation in natural language, and emotionally sensitive responses are provided, thereby reducing staff stress and enabling efficient and smooth operation.

[0670] "Input means" refers to a device or interface for a user to input instructions in natural language.

[0671] "Analysis means" refers to a device or program for analyzing the user's input and emotional state.

[0672] "Generation means" refers to a device or program for generating a script or command based on the analyzed content and emotional state.

[0673] "Presentation means" refers to a device or interface for presenting a generated script or command and an emotionally sensitive response message to the user.

[0674] An "emotion recognition engine" is a program or device that analyzes a user's facial expressions and voice tone to identify their emotional state.

[0675] A "natural language processing engine" is a program or device used to analyze the content of natural language input by a user.

[0676] A "server" is a central processing unit that executes analysis and generation methods and provides the results to the user.

[0677] A "script" is a series of instructions generated to automate a specific task.

[0678] A "command" is an instruction issued to perform a specific operation or task.

[0679] A "response message" is additional information or a warning message provided with consideration for the user's emotional state.

[0680] This invention is a system in which a user inputs instructions in natural language, generates a script or command based on the analyzed content, and then presents the generated result. The invention incorporates an emotion recognition engine that recognizes the user's emotions, thereby enabling it to provide appropriate responses according to the user's emotions.

[0681] System Overview

[0682] This system includes the following main components:

[0683] 1. Input method: The user inputs information using natural language. This includes smartphone apps and factory robot interfaces.

[0684] 2. Analysis Method: The user's input and emotional state are analyzed. This includes the Google Cloud Natural Language API and the emotion recognition engine.

[0685] 3. Generation method: Generates a script or command based on the analyzed content and emotional state. OpenAI's generative AI model is an example of this.

[0686] 4. Presentation method: The generated script or command and an emotionally sensitive response message are presented to the user. This includes smartphone apps and robot display devices.

[0687] Program processing

[0688] The server first receives natural language instructions entered by the user. Next, it uses an emotion recognition engine to analyze the user's facial expressions and tone of voice to identify their emotional state. The analysis method uses the Google Cloud Natural Language API to analyze and understand the user's instructions.

[0689] In the generation process, OpenAI's generation AI model generates appropriate scripts or commands based on the analysis results and emotional state. Simultaneously, response messages corresponding to the user's emotional state are also generated. For example, if the user is feeling stressed, a reassuring message will be generated.

[0690] The generated script or command and response messages are presented to the user via a smartphone app or the robot's display device. The user can then proceed with their task using the presented script or command.

[0691] Specific examples and prompt statements

[0692] For example, consider a scenario where a user instructs a factory line to stop.

[0693] User instruction: "Stop factory line 1."

[0694] Example of a prompt:

[0695] Effect: Stop factory line 1

[0696] Emotion Score: -0.5

[0697] Generate the appropriate command and supportive message.

[0698] In this case, the server generates a response message along with the command "stop line 1" that reads, "The operation is complete. Please let us know if there is anything else we can help you with."

[0699] This system allows users to intuitively control robots using natural language without having to input complex commands, enabling them to work efficiently and without experiencing stress during the process.

[0700] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0701] Step 1:

[0702] The user inputs instructions. The user inputs instructions in natural language using a smartphone app or the factory robot's interface. The entered natural language data is sent to the server.

[0703] Step 2:

[0704] The server receives the input. The server receives natural language data sent by the user and stores its contents. The input is natural language text data. The output is the storage and management of this text data.

[0705] Step 3:

[0706] The system performs emotion analysis using an emotion recognition engine. The server passes the input natural language data to the Google Cloud Natural Language API, which analyzes the user's emotional state from voice tone and facial expression data (which may be obtained via webcam or microphone). The input consists of natural language text data and voice / facial expression data, and the output is an emotional state score.

[0707] Step 4:

[0708] This system performs natural language processing. The server passes input data to the Google Cloud Natural Language API for language analysis. The input is natural language text data, and the output is the analyzed content (text structure and semantic information as a result of the analysis).

[0709] Step 5:

[0710] The server generates scripts and response messages. It passes the analysis results and emotional state scores to OpenAI's generative AI model, which then generates appropriate scripts or commands and emotionally sensitive response messages. The input is the analysis results and emotional state scores, and the output is the generated scripts or commands and response messages.

[0711] Step 6:

[0712] The generated results are presented. The server sends the generated script or command and response message to the user's terminal. The user receives the presented results through a smartphone app or a robot's display device. The input is the generated script or command and response message, and the output is the display on the user interface.

[0713] Step 7:

[0714] User verification and execution. The user reviews the presented script or command and response messages and uses them to proceed with the task. The user provides further input as needed, and the process continues.

[0715] These steps enable users to easily generate and execute scripts or commands based on emotionally sensitive natural language instructions.

[0716] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0717] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0718] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0719] [Third Embodiment]

[0720] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0721] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0722] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0723] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0724] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0725] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0726] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0727] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0728] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0729] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0730] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0731] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0732] Overall Overview

[0733] The present invention provides a system that automatically generates and presents necessary scripts and commands simply by having the user input instructions in natural language. This system supports users in carrying out their tasks quickly and accurately.

[0734] System Configuration

[0735] This system includes the following main components:

[0736] 1. Input method: The user inputs using natural language.

[0737] 2. Analysis method: Analyzes the user's input.

[0738] 3. Generation method: Generate a script or command based on the analyzed content.

[0739] 4. Presentation method: The generated script or command is presented to the user.

[0740] Details of the example

[0741] 1. User input

[0742] The user inputs instructions into the terminal using natural language. For example, they might input, "I want to generate a command to connect to a specified server and check disk usage."

[0743] 2. Natural language analysis

[0744] The terminal sends user input to the server. The server analyzes the input using a natural language processing engine. The processing engine understands what the user wants to achieve and extracts information to generate appropriate scripts or commands.

[0745] 3. Script / command generation

[0746] Based on the analyzed information, the server generates appropriate scripts and commands. For example, the server generates the Unix command "df -h" to check disk usage.

[0747] 4. Presentation of results

[0748] The server generates scripts and commands, which are then sent back to the terminal. The terminal then presents these to the user. The user can review the generated commands and use them.

[0749] Specific example

[0750] Example 1: Checking disk usage

[0751] User input: The user instructs the terminal to "generate a command to check the server's disk usage."

[0752] Server analysis and generation: The server analyzes the user's instructions and generates the disk usage command "df -h".

[0753] Prompt: The terminal prompts the user to execute "df -h". The user uses the suggested command to check the server's disk usage.

[0754] Example 2: Retrieving a file list

[0755] User input: The user instructs the terminal to "generate a command to retrieve a list of files in the specified directory."

[0756] Server analysis and generation: The server analyzes the data using commands like "ls -la / path / to / directory" and generates the appropriate commands.

[0757] Presentation: The terminal presents the user with the command "ls -la / path / to / directory", and the user executes the command to obtain a list of files.

[0758] Operation flow and effects

[0759] This system allows users to quickly and easily generate appropriate scripts and commands, enabling them to work efficiently. Because users can generate complex commands simply by inputting instructions in natural language, the system can be used without problems even by those lacking technical knowledge. Furthermore, the generated scripts and commands can be visually reviewed, allowing users to proceed with confidence.

[0760] The following describes the processing flow.

[0761] Step 1:

[0762] The user enters instructions into the terminal using natural language. For example, they might enter, "I want to generate a command to connect to a specified server and check disk usage."

[0763] Step 2:

[0764] The terminal receives user input and sends that input to the server. The data sent to the server is natural language text data.

[0765] Step 3:

[0766] The server sends the data to a natural language processing engine to analyze the user's input. The natural language processing engine analyzes the input sentence and extracts information to understand the user's intent.

[0767] Step 4:

[0768] The natural language processing engine sends the analysis results back to the server. The analysis results include the user's desired objective (for example, "check disk usage").

[0769] Step 5:

[0770] The server generates an appropriate script or command based on the analysis results. For example, it generates the command "df -h" to check disk usage.

[0771] Step 6:

[0772] The server sends the generated script or command back to the terminal. The terminal receives this data.

[0773] Step 7:

[0774] The terminal displays the received script or command to the user. The user confirms the displayed content on the terminal screen.

[0775] Step 8:

[0776] If a user needs to execute a presented script or command, they send the command to the server via the terminal. The server executes the received command and returns the result to the user.

[0777] Step 9:

[0778] The terminal displays the execution results of commands received from the server to the user. The user can then obtain the necessary information.

[0779] Through the above processing steps, users can efficiently perform a series of operations, from natural language input to the generation and execution of appropriate scripts and commands. This system makes it easy to generate and execute complex commands, even for those lacking technical knowledge.

[0780] (Example 1)

[0781] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0782] In conventional systems, users needed knowledge of commands and scripts to perform the desired operations, making them difficult for non-specialized users. Furthermore, generating scripts and commands had to be done manually, which was time-consuming and laborious. Therefore, there was a need for a system that would allow users to easily generate commands using natural language.

[0783] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0784] In this invention, the server includes an input means for the user to input in natural language, a transmission means for the terminal to send the user's input to the server, an analysis means for the server to analyze the user's input, a generation means to generate a script or command based on the analyzed content, and a presentation means to return the script or command generated from the server to the terminal and present it to the user. This makes it possible for the user to easily generate the desired script or command in natural language and to execute operations quickly and accurately.

[0785] "Natural language" refers to the language that humans use on a daily basis, and is not a specific programming language or specialized terminology, but rather a general language.

[0786] "Input method" refers to a device or interface that allows a user to input information or instructions in natural language.

[0787] "Transmission means" refers to the communication functions and protocols that a terminal uses to send user input to a server.

[0788] "Analysis means" refers to algorithms and engines that semantically understand the natural language input received by the server and perform appropriate processing.

[0789] A "natural language processing engine" refers to software or a system that analyzes input natural language and understands its meaning and intent.

[0790] "Generation means" refers to algorithms or engines used to generate specific scripts or commands based on the analyzed content.

[0791] A "template database" refers to a set of predefined templates that a generation tool references when generating scripts or commands.

[0792] "Presentation means" refers to devices or interfaces used to display generated scripts or commands to the user.

[0793] A "terminal" refers to a computer or mobile device used by a user to input information and receive generated scripts or commands.

[0794] A "server" refers to a computer system equipped with analysis and generation capabilities, designed to process input from terminals.

[0795] Overall Overview

[0796] The present invention provides a system that automatically generates and presents necessary scripts and commands simply by having the user input instructions in natural language. This allows users to perform operations efficiently even without specialized knowledge. The system provides support for users to carry out their tasks quickly and accurately.

[0797] System Configuration

[0798] This system includes the following main components:

[0799] 1. Input method: The user inputs using natural language.

[0800] 2. Transmission method: The terminal sends the user's input to the server.

[0801] 3. Analysis method: The server analyzes the user's input.

[0802] 4. Generation method: Generate a script or command based on the analyzed content.

[0803] 5. Presentation method: The generated script or command is presented to the user.

[0804] Details of the example

[0805] 1. User input

[0806] The user inputs instructions into the terminal using natural language. For example, they might input, "Please generate a command to connect to the specified server and check disk usage." An example of a specific prompt is as follows:

[0807] Example of a prompt:

[0808] "Please generate a command to connect to the specified server and check disk usage."

[0809] "Generate a command to retrieve a list of files in a specified directory."

[0810] 2. Natural language analysis

[0811] The terminal sends the user's input to the server. The server analyzes the input using a natural language processing engine (e.g., "GPT-4"). The natural language processing engine syntactically analyzes the input text, understands its intent, and then extracts the necessary information.

[0812] 3. Script / command generation

[0813] Based on the analyzed data, the server generates appropriate scripts and commands. In this process, the server uses a command generation engine (e.g., a "shell script generation library") to generate specific commands by supplementing templates obtained from a template database. For example, to check disk usage, the Unix command "df -h" is generated.

[0814] 4. Presentation of results

[0815] The generated scripts and commands are sent back from the server to the terminal. The terminal visually presents them to the user, who can then review and use them. For example, the generated command "df -h" is displayed on the terminal screen and executed by the user.

[0816] Specific example

[0817] Example 1: Checking disk usage

[0818] User input: "Generate a command to check the server's disk usage."

[0819] Server analysis and generation: The server analyzes the input and generates the command "df -h".

[0820] Instructions: The terminal will display "df -h" to the user, who will use this to check disk usage.

[0821] Example 2: Retrieving a file list

[0822] User input: "Generate a command to get a list of files in the specified directory."

[0823] Server analysis and generation: The server generates the command "ls -la / path / to / directory".

[0824] Instructions: The terminal displays "ls -la / path / to / directory" to the user, who then uses this to obtain a list of files.

[0825] Operation flow and effects

[0826] This system allows users to quickly and easily generate appropriate scripts and commands. Even without technical knowledge, complex commands can be generated simply by inputting instructions in natural language. Therefore, users can perform their tasks efficiently, and because they can visually confirm the generated scripts and commands, they can proceed with confidence.

[0827] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0828] Step 1:

[0829] The user inputs instructions into the terminal using natural language. Specifically, the user inputs text data such as, "Please generate a command to connect to the specified server and check disk usage." This input becomes the input data required to proceed to the next step.

[0830] Step 2:

[0831] The terminal sends the user's input to the server. The terminal sends the entered text as a string to the server, which is then processed by the parsing engine. In this step, the input text is transferred from the terminal to the server.

[0832] Step 3:

[0833] The server analyzes the user's input. The server uses a natural language processing engine to analyze the input and understand the user's intent. In this process, the input data (natural language text) is syntactically analyzed and semantically. The analysis engine extracts the intent "check disk usage" from the input text and generates analysis results to proceed to the next step.

[0834] Step 4:

[0835] The server generates a script or command based on the analyzed data. Here, the server uses a command generation engine to retrieve the appropriate template from the template database and generate a specific command. For example, if the analysis result is "check disk usage", the command "df -h" will be generated. Based on the analysis result (parsed data), the generation engine processes the data and generates a specific command as output.

[0836] Step 5:

[0837] The server sends the generated scripts and commands back to the terminal. Specifically, the server sends the command "df -h" generated by the generation engine to the terminal, and a process of receiving it takes place.

[0838] Step 6:

[0839] The terminal presents the generated script or command to the user. The terminal displays the received command in the user interface, allowing the user to review and use it. For example, the command "df -h" might appear on the screen, and the user can copy and execute that command.

[0840] (Application Example 1)

[0841] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0842] Traditionally, supporting the operation of factory robots often requires advanced technical knowledge and specialized skills to properly execute complex operations and maintenance instructions. This presents challenges, particularly in responding quickly and accurately to emergencies and non-routine tasks. It also increases the risk of operational errors and work delays.

[0843] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0844] In this invention, the server includes an input means for the user to input information in natural language, an analysis means for analyzing the user's input, a generation means for generating a script or command based on the analyzed content, a presentation means for presenting the generated script or command to the user, and an execution means for applying the generated script or command to a machine. This enables the user to quickly and accurately operate and maintain factory robots without requiring specialized knowledge.

[0845] "Input means" refers to functions and devices that allow users to input commands and instructions in natural language.

[0846] "Analysis means" refers to functions and devices for analyzing the content of natural language obtained from input means and understanding the user's intent.

[0847] "Generation means" refers to a function and device for generating an appropriate script or command based on the content analyzed by the analysis means.

[0848] "Presentation means" refers to functions and devices for presenting scripts or commands generated by the generation means to the user visually or by other means.

[0849] "Execution means" refers to functions and devices for actually applying the script or command presented by the presentation means to a machine and performing operations or tasks.

[0850] A "generative AI model" is an artificial intelligence model that analyzes natural language and generates appropriate scripts or commands based on the analysis results.

[0851] This invention provides a system that streamlines the operation of factory robots, enabling users to generate and execute appropriate operation scripts and commands simply by inputting instructions in natural language. This system includes the following main components:

[0852] An "input method" refers to a function or device that allows a user to input commands or instructions in natural language. Specifically, this includes devices such as smartphones, tablets, and factory terminals. The user inputs instructions to the device in natural language, such as "Return the robot arm to the home position."

[0853] "Analysis means" refers to functions and devices for analyzing the content of natural language obtained from input means and understanding the user's intent. This system uses a generative AI model (e.g., OpenAI GPT-3) as the analysis engine. This model analyzes the input natural language using advanced algorithms and extracts the intended operation.

[0854] The "generation means" refers to a function and device for generating an appropriate script or command based on the content analyzed by the analysis means. If the analyzed content is an instruction to "return the robot arm to the home position," the generation AI model will generate a script such as "moveToHomePosition();". The generation means is mainly executed on a server, and the generated script is processed within the server.

[0855] "Presentation means" refers to functions and devices for presenting the script or command generated by the generation means to the user visually or by other means. Specifically, it displays the generated script on the screen of a smartphone or tablet. The user can then review its contents and decide whether or not to execute it.

[0856] "Execution means" refers to the functions and devices that actually apply the script or command presented by the presentation means to the machine and perform the operation or task. Once the user confirms the displayed script and approves its execution, the script is sent to the factory robot and the actual operation begins.

[0857] Specific example

[0858] For example, if a user enters "Get the current position of the robot arm" into a tablet, the system's analysis mechanism analyzes the input and uses a generation AI model to generate the command "getCurrentPosition();". The generated command is then displayed on the tablet screen for the user to confirm. After the user confirms, the command is sent to the robot, which obtains its current position and returns that information.

[0859] Example of a prompt

[0860] The following are examples of specific prompt statements used for generative AI models.

[0861] Based on the following natural language instructions, generate a script to operate a factory robot:

[0862] Instructions: Return the robot arm to its home position.

[0863] script:

[0864] In this way, by allowing users to input instructions in natural language and automatically generating appropriate scripts and commands based on those instructions, the operation of factory robots is simplified, and efficient work is achieved.

[0865] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0866] Step 1:

[0867] The user inputs a natural language instruction, such as "Return the robot arm to the home position," into an input device such as a tablet or smartphone. This input information is then transmitted to the system's input method.

[0868] Input: User's natural language instructions

[0869] Output: Instruction data to the input means

[0870] Step 2:

[0871] Natural language instructions received through the input device are sent to the server. The server's analysis device inputs these instructions into a generative AI model (e.g., OpenAI GPT-3) and analyzes the content. The generative AI model analyzes the input natural language and processes the data to understand the user's intent.

[0872] Input: Natural language instruction data

[0873] Output: Analyzed instruction data

[0874] Step 3:

[0875] The generating AI model generates an appropriate script or command based on the analyzed instruction data. For example, in response to the instruction "Return the robot arm to the home position," the script "moveToHomePosition();" is generated. This generated script is then processed by the generation mechanism.

[0876] Input: Analyzed instruction data

[0877] Output: Generated script or command

[0878] Step 4:

[0879] The generated script or command is visually displayed to the user through a presentation mechanism. The user reviews the script on their tablet or smartphone screen. At this point, the user reviews the content of the generated command and decides whether or not to execute it.

[0880] Input: Generated script or command

[0881] Output: Displayed script

[0882] Step 5:

[0883] After the user reviews the presented script and presses the execute button, the script is sent to the factory robot via the server. This causes the robot to begin its actual operation via the execution mechanism. For example, the robot arm might perform an action to return to its home position.

[0884] Input: User verification and execution command

[0885] Output: Actual operation by the robot

[0886] In this process, users can quickly and accurately operate factory robots simply by inputting instructions in natural language. By coordinating the generative AI model, presentation method, and execution method, it is possible to significantly reduce the burden on the user and achieve efficient work.

[0887] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0888] Overall Overview

[0889] The present invention provides a system that automatically generates and presents necessary scripts and commands simply by having the user input instructions in natural language. Furthermore, the present invention incorporates an emotion engine that recognizes the user's emotions, enabling it to provide appropriate responses according to the user's feelings. This system provides support for users to carry out their tasks quickly and accurately.

[0890] System Configuration

[0891] This system includes the following main components:

[0892] 1. Input method: The user inputs using natural language.

[0893] 2. Analysis method: Analyzes the user's input.

[0894] 3. Generation method: Generate a script or command based on the analyzed content.

[0895] 4. Presentation method: The generated script or command is presented to the user.

[0896] 5. Emotion Engine: Recognizes the user's emotions and provides feedback to analysis and generation methods.

[0897] Details of the example

[0898] 1. User Input and Sentiment Recognition

[0899] The user inputs instructions to the terminal using natural language. For example, if the user inputs, "I want to generate a command to connect to a specified server and check disk usage," the emotion engine simultaneously recognizes the user's emotions (e.g., frustration or tension) from the user's input, voice, facial expressions, and gestures.

[0900] 2. Analysis of Natural Language and Sentiment

[0901] The terminal sends user input and emotional information to the server. The server uses a natural language processing engine to analyze the input and an emotional processing engine to analyze the user's emotions. The analysis results include the user's desired objective (e.g., "check disk usage") and the user's emotional state.

[0902] 3. Script / command generation and emotion reflection

[0903] Based on the analyzed data, the server generates appropriate scripts and commands. For example, it generates the Unix command "df -h" to check disk usage. Simultaneously, the emotion engine generates additional information and warning messages to alleviate user frustration and tension (e.g., "Executing this command will show you your disk usage").

[0904] 4. Presentation of results and implementation

[0905] The server generates scripts, commands, and emotionally sensitive messages, which are then sent back to the terminal. The terminal presents these to the user. The user can then review the generated commands and additional information and use them.

[0906] Specific example

[0907] Example 1: Checking disk usage

[0908] User input: The user instructs the terminal to "generate a command to check the server's disk usage."

[0909] Server analysis and generation: The server analyzes the user's instructions and generates the disk usage command "df -h". The emotion engine detects the user's tension and generates an additional message: "Executing this command will show you your disk usage."

[0910] Prompt: The terminal prompts the user with the command "df -h" and an additional message. The user uses the provided command to check the server's disk usage.

[0911] Example 2: Retrieving a file list

[0912] User input: The user instructs the terminal to "generate a command to retrieve a list of files in the specified directory."

[0913] Server analysis and generation: The server analyzes the data, such as "ls -la / path / to / directory", and generates the appropriate command. The sentiment engine detects user frustration and generates an additional message such as, "Executing this command will show you a list of files in the specified directory."

[0914] Presentation: The terminal presents the user with the additional message "ls -la / path / to / directory", and the user executes the command to obtain a list of files.

[0915] Operation flow and effects

[0916] This system allows users to quickly and easily generate appropriate scripts and commands, and receive responses that take their emotions into consideration thanks to its emotion engine. Users not only input instructions in natural language, but the system also enhances the user experience by providing additional information and warning messages based on the user's emotions. Even those lacking technical knowledge can use the system without difficulty, allowing for confident operation.

[0917] The following describes the processing flow.

[0918] Step 1:

[0919] The user enters instructions into the terminal using natural language. For example, they might enter, "I want to generate a command to connect to a specified server and check disk usage."

[0920] Step 2:

[0921] The emotion engine recognizes the user's emotions from their input, voice, facial expressions, and other factors. For example, it can detect user frustration based on the tone of their voice and input speed.

[0922] Step 3:

[0923] The device sends user instructions in natural language and emotional information to the server. The data sent to the server consists of natural language text data and emotional information generated by an emotion engine.

[0924] Step 4:

[0925] The server uses a natural language processing engine to analyze the user's instructions. The processing engine identifies the user's desired objective (for example, "check disk usage").

[0926] Step 5:

[0927] The emotion engine analyzes the transmitted emotion information to identify the user's emotional state. For example, it determines whether the user is feeling "stressed" or "frustrated."

[0928] Step 6:

[0929] The server generates an appropriate script or command based on the analyzed information. For example, it might generate the Unix command "df -h" to check disk usage.

[0930] Step 7:

[0931] The server generates additional messages tailored to the user's emotions based on the analysis results of the emotion engine. For example, if the user is feeling frustrated, it will generate encouraging and reassuring messages such as, "Running this command will show you your disk usage."

[0932] Step 8:

[0933] The server sends the generated script or command and any additional messages to the terminal.

[0934] Step 9:

[0935] The terminal displays the received script or command and any additional messages to the user. The user reviews the displayed content on the terminal screen.

[0936] Step 10:

[0937] If a user needs to execute a presented script or command, they send the command to the server via the terminal. The server executes the received command and returns the result to the user.

[0938] Step 11:

[0939] The terminal displays the execution results of commands received from the server to the user. The user can then obtain the necessary information.

[0940] Through the processing steps described above, users can efficiently perform a series of operations, from natural language input to the generation of appropriate scripts and commands, to emotionally sensitive responses and execution. This system allows users to carry out their work quickly and easily, and the emotional engine provides a less stressful user experience.

[0941] (Example 2)

[0942] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0943] Current scripting and command generation systems require users to possess a high level of technical knowledge and have limited understanding of natural language instructions, thus limiting user convenience. Furthermore, they fail to provide responses that consider the user's emotional state, resulting in a lack of appropriate support, especially for users experiencing frustration or anxiety. Therefore, there is a need for a system that can quickly and accurately understand user instructions and respond with consideration for the user's emotions.

[0944] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0945] In this invention, the server includes an input means for the user to input in natural language, an analysis means for analyzing the user's input and emotional information, a generation means for generating a script or command based on the analyzed content and emotional information, and a presentation means for presenting the generated script or command and additional emotionally sensitive information to the user. As a result, the user can quickly and accurately generate the necessary scripts and commands simply by inputting instructions in natural language, and the system provides responses that are sensitive to the user's emotions, thereby improving the user experience.

[0946] "Input means" refers to a device or interface for a user to input instructions in natural language.

[0947] "Analysis means" refers to a device or software for analyzing user input and emotional information.

[0948] "Generation means" refers to a device or software for generating scripts or commands based on analyzed content and emotional information.

[0949] "Presentation means" refers to a device or interface for presenting a generated script or command and additional information that takes emotions into consideration to the user.

[0950] A "natural language processing engine" is software or an algorithm that analyzes natural language input by a user to understand its intent and purpose.

[0951] An "emotion analysis engine" is software or an algorithm that analyzes a user's emotional state from inputs such as voice, facial expressions, and gestures.

[0952] A "script" is a set of commands or program code generated to automate a specific task or operation.

[0953] A "command" is an instruction generated to tell a system to perform a specific operation or task.

[0954] "Additional information" refers to supplementary messages or explanations provided alongside generated scripts or commands, which are designed with the user's feelings in mind.

[0955] Overall Overview

[0956] The system of this invention automatically generates and presents necessary scripts and commands simply by the user inputting in natural language. Furthermore, the invention incorporates an emotion engine that recognizes the user's emotions, enabling it to provide appropriate responses according to the user's feelings. This system provides support that allows users to carry out their tasks quickly and accurately, even if they lack technical knowledge.

[0957] System Configuration

[0958] This system includes the following main components:

[0959] 1. Input means: A device or interface in which the user inputs information in natural language.

[0960] 2. Analysis means: A device or software that analyzes user input and emotional information.

[0961] 3. Generation means: A device or software that generates a script or command based on the analyzed content and emotional information.

[0962] 4. Presentation means: A device or interface that presents the generated script or command and additional information that takes emotions into consideration to the user.

[0963] Hardware and software to be used

[0964] The following hardware and software will be used for the specific implementation of the system:

[0965] Natural language processing engine: For example, software such as NLTK or OpenAI's GPT-4 can be used.

[0966] Sentiment analysis engine: For example, software such as Google's Dialogflow can be used.

[0967] Device: A computer, smartphone, or other device that the user directly operates.

[0968] Server: A remote server for executing analysis or generation methods.

[0969] Specific examples of how the system works

[0970] User input and emotion recognition

[0971] The user inputs instructions to the terminal using natural language. For example, they might input, "I want to generate a command to connect to the specified server and check disk usage." The terminal also recognizes the user's emotions (e.g., frustration or tension) from their input, voice, facial expressions, and gestures.

[0972] Analysis of natural language and emotions

[0973] The terminal sends user input and emotional information to the server. The server uses a natural language processing engine to analyze the input and an emotional processing engine to analyze the user's emotions. The analysis results include the user's desired objective (e.g., "check disk usage") and the user's emotional state.

[0974] Script / command generation and emotion reflection

[0975] Based on the analyzed data, the server generates appropriate scripts and commands. For example, it generates the Unix command "df -h" to check disk usage. At the same time, the emotion engine generates additional information to alleviate user frustration and tension (for example, "Executing this command will show you your disk usage").

[0976] Presentation of results and execution

[0977] The server generates scripts, commands, and emotionally sensitive messages, which are then sent back to the terminal. The terminal presents these to the user. The user can then review the generated commands and additional information and use them.

[0978] Example of a prompt

[0979] Here are some examples of specific prompt messages:

[0980] "Please generate Unix commands to connect to the specified server and check disk usage. Considering the user's anxiety, please also include an explanation of the execution results."

[0981] "Generate a Unix command to retrieve a list of files in a specified directory. Include a message explaining how this command can be helpful if the user is experiencing frustration."

[0982] In this way, the present invention realizes a system that flexibly and quickly analyzes the user's natural language input and provides an appropriate response that takes emotions into consideration.

[0983] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0984] Step 1:

[0985] The user enters information in natural language.

[0986] The user inputs instructions into the terminal using natural language. For example, they might input, "Generate a command to check the server's disk usage." This input is saved to the terminal as text data.

[0987] Step 2:

[0988] The device retrieves input content and emotional information.

[0989] The device acquires not only the user's input but also emotional information such as the user's voice, facial expressions, and gestures. As a result, the input data includes both the user's text data and emotional data.

[0990] Step 3:

[0991] The terminal sends the input data to the server.

[0992] The device sends collected text and sentiment data to the server. This data is packaged in JSON format or similar and transferred to the server over the network.

[0993] Step 4:

[0994] The server analyzes natural language.

[0995] The server analyzes the received text data using a natural language processing engine (such as NLTK or OpenAI's GPT-4). Specifically, it analyzes the user's instructions and understands the objective, which is "checking disk usage." This analysis result is then generated as intermediate data.

[0996] Step 5:

[0997] The server analyzes emotional information.

[0998] The server uses an emotion analysis engine (Google's Dialogflow) to analyze emotional data. For example, it can identify if a user is feeling nervous based on their voice tone and facial expressions. This analysis result is also generated as intermediate data.

[0999] Step 6:

[1000] The server generates commands or scripts.

[1001] The server integrates natural language processing results and sentiment analysis results to generate appropriate scripts and commands. For example, it generates the Unix command "df -h" to check disk usage. This generated command is then output.

[1002] Step 7:

[1003] The server generates additional information that takes emotions into consideration.

[1004] Based on the sentiment analysis results, the server generates additional information and explanations that take the user's emotions into consideration. For example, it might generate a message such as, "Executing this command will show you your disk usage." The command message with this additional information is then output.

[1005] Step 8:

[1006] The server sends the generated data to the terminal.

[1007] The server sends the generated scripts, commands, and additional information to the terminal. This data is also structured in JSON format or similar and transferred to the terminal over the network.

[1008] Step 9:

[1009] The device presents to the user

[1010] The terminal presents the received scripts, commands, and any additional messages to the user. The user reviews them and, if necessary, executes the commands directly on the terminal.

[1011] (Application Example 2)

[1012] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1013] Conventional factory robot systems require staff to perform complex coding and command input when issuing operating instructions to the robots, thus demanding specialized knowledge. Furthermore, a lack of consideration for staff emotional states often leads to stress and frustration. There is a need to solve these problems and make factory robot operation more efficient and user-friendly.

[1014] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an input means for the user to input in natural language, an analysis means for analyzing the user's input content and emotional state, a generation means for generating a script or command based on the analyzed content and emotional state, and a presentation means for presenting the generated script or command and an emotionally sensitive response message to the user. As a result, factory staff can easily instruct robot operation in natural language, and emotionally sensitive responses are provided, thereby reducing staff stress and enabling efficient and smooth operation.

[1015] "Input means" refers to a device or interface for a user to input instructions in natural language.

[1016] "Analysis means" refers to a device or program for analyzing the user's input and emotional state.

[1017] "Generation means" refers to a device or program for generating a script or command based on the analyzed content and emotional state.

[1018] "Presentation means" refers to a device or interface for presenting a generated script or command and an emotionally sensitive response message to the user.

[1019] An "emotion recognition engine" is a program or device that analyzes a user's facial expressions and voice tone to identify their emotional state.

[1020] A "natural language processing engine" is a program or device used to analyze the content of natural language input by a user.

[1021] A "server" is a central processing unit that executes analysis and generation methods and provides the results to the user.

[1022] A "script" is a series of instructions generated to automate a specific task.

[1023] A "command" is an instruction issued to perform a specific operation or task.

[1024] A "response message" is additional information or a warning message provided with consideration for the user's emotional state.

[1025] This invention is a system in which a user inputs instructions in natural language, generates a script or command based on the analyzed content, and then presents the generated result. The invention incorporates an emotion recognition engine that recognizes the user's emotions, thereby enabling it to provide appropriate responses according to the user's emotions.

[1026] System Overview

[1027] This system includes the following main components:

[1028] 1. Input method: The user inputs information using natural language. This includes smartphone apps and factory robot interfaces.

[1029] 2. Analysis Method: The user's input and emotional state are analyzed. This includes the Google Cloud Natural Language API and the emotion recognition engine.

[1030] 3. Generation method: Generates a script or command based on the analyzed content and emotional state. OpenAI's generative AI model is an example of this.

[1031] 4. Presentation method: The generated script or command and an emotionally sensitive response message are presented to the user. This includes smartphone apps and robot display devices.

[1032] Program processing

[1033] The server first receives natural language instructions entered by the user. Next, it uses an emotion recognition engine to analyze the user's facial expressions and tone of voice to identify their emotional state. The analysis method uses the Google Cloud Natural Language API to analyze and understand the user's instructions.

[1034] In the generation process, OpenAI's generation AI model generates appropriate scripts or commands based on the analysis results and emotional state. Simultaneously, response messages corresponding to the user's emotional state are also generated. For example, if the user is feeling stressed, a reassuring message will be generated.

[1035] The generated script or command and response messages are presented to the user via a smartphone app or the robot's display device. The user can then proceed with their task using the presented script or command.

[1036] Specific examples and prompt statements

[1037] For example, consider a scenario where a user instructs a factory line to stop.

[1038] User instruction: "Stop factory line 1."

[1039] Example of a prompt:

[1040] Effect: Stop factory line 1

[1041] Emotion Score: -0.5

[1042] Generate the appropriate command and supportive message.

[1043] In this case, the server generates a response message along with the command "stop line 1" that reads, "The operation is complete. Please let us know if there is anything else we can help you with."

[1044] This system allows users to intuitively control robots using natural language without having to input complex commands, enabling them to work efficiently and without experiencing stress during the process.

[1045] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1046] Step 1:

[1047] The user inputs instructions. The user inputs instructions in natural language using a smartphone app or the factory robot's interface. The entered natural language data is sent to the server.

[1048] Step 2:

[1049] The server receives the input. The server receives natural language data sent by the user and stores its contents. The input is natural language text data. The output is the storage and management of this text data.

[1050] Step 3:

[1051] The system performs emotion analysis using an emotion recognition engine. The server passes the input natural language data to the Google Cloud Natural Language API, which analyzes the user's emotional state from voice tone and facial expression data (which may be obtained via webcam or microphone). The input consists of natural language text data and voice / facial expression data, and the output is an emotional state score.

[1052] Step 4:

[1053] This system performs natural language processing. The server passes input data to the Google Cloud Natural Language API for language analysis. The input is natural language text data, and the output is the analyzed content (text structure and semantic information as a result of the analysis).

[1054] Step 5:

[1055] The server generates scripts and response messages. It passes the analysis results and emotional state scores to OpenAI's generative AI model, which then generates appropriate scripts or commands and emotionally sensitive response messages. The input is the analysis results and emotional state scores, and the output is the generated scripts or commands and response messages.

[1056] Step 6:

[1057] The generated results are presented. The server sends the generated script or command and response message to the user's terminal. The user receives the presented results through a smartphone app or a robot's display device. The input is the generated script or command and response message, and the output is the display on the user interface.

[1058] Step 7:

[1059] User verification and execution. The user reviews the presented script or command and response messages and uses them to proceed with the task. The user provides further input as needed, and the process continues.

[1060] These steps enable users to easily generate and execute scripts or commands based on emotionally sensitive natural language instructions.

[1061] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1062] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1063] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1064] [Fourth Embodiment]

[1065] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1066] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1067] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1068] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1069] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1070] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1071] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1072] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1073] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1074] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1075] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1076] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1077] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1078] Overall Overview

[1079] The present invention provides a system that automatically generates and presents necessary scripts and commands simply by having the user input instructions in natural language. This system supports users in carrying out their tasks quickly and accurately.

[1080] System Configuration

[1081] This system includes the following main components:

[1082] 1. Input method: The user inputs using natural language.

[1083] 2. Analysis method: Analyzes the user's input.

[1084] 3. Generation method: Generate a script or command based on the analyzed content.

[1085] 4. Presentation method: The generated script or command is presented to the user.

[1086] Details of the example

[1087] 1. User input

[1088] The user inputs instructions into the terminal using natural language. For example, they might input, "I want to generate a command to connect to a specified server and check disk usage."

[1089] 2. Natural language analysis

[1090] The terminal sends user input to the server. The server analyzes the input using a natural language processing engine. The processing engine understands what the user wants to achieve and extracts information to generate appropriate scripts or commands.

[1091] 3. Script / command generation

[1092] Based on the analyzed information, the server generates appropriate scripts and commands. For example, the server generates the Unix command "df -h" to check disk usage.

[1093] 4. Presentation of results

[1094] The server generates scripts and commands, which are then sent back to the terminal. The terminal then presents these to the user. The user can review the generated commands and use them.

[1095] Specific example

[1096] Example 1: Checking disk usage

[1097] User input: The user instructs the terminal to "generate a command to check the server's disk usage."

[1098] Server analysis and generation: The server analyzes the user's instructions and generates the disk usage command "df -h".

[1099] Prompt: The terminal prompts the user to execute "df -h". The user uses the suggested command to check the server's disk usage.

[1100] Example 2: Retrieving a file list

[1101] User input: The user instructs the terminal to "generate a command to retrieve a list of files in the specified directory."

[1102] Server analysis and generation: The server analyzes the data using commands like "ls -la / path / to / directory" and generates the appropriate commands.

[1103] Presentation: The terminal presents the user with the command "ls -la / path / to / directory", and the user executes the command to obtain a list of files.

[1104] Operation flow and effects

[1105] This system allows users to quickly and easily generate appropriate scripts and commands, enabling them to work efficiently. Because users can generate complex commands simply by inputting instructions in natural language, the system can be used without problems even by those lacking technical knowledge. Furthermore, the generated scripts and commands can be visually reviewed, allowing users to proceed with confidence.

[1106] The following describes the processing flow.

[1107] Step 1:

[1108] The user enters instructions into the terminal using natural language. For example, they might enter, "I want to generate a command to connect to a specified server and check disk usage."

[1109] Step 2:

[1110] The terminal receives user input and sends that input to the server. The data sent to the server is natural language text data.

[1111] Step 3:

[1112] The server sends the data to a natural language processing engine to analyze the user's input. The natural language processing engine analyzes the input sentence and extracts information to understand the user's intent.

[1113] Step 4:

[1114] The natural language processing engine sends the analysis results back to the server. The analysis results include the user's desired objective (for example, "check disk usage").

[1115] Step 5:

[1116] The server generates an appropriate script or command based on the analysis results. For example, it generates the command "df -h" to check disk usage.

[1117] Step 6:

[1118] The server sends the generated script or command back to the terminal. The terminal receives this data.

[1119] Step 7:

[1120] The terminal displays the received script or command to the user. The user confirms the displayed content on the terminal screen.

[1121] Step 8:

[1122] If a user needs to execute a presented script or command, they send the command to the server via the terminal. The server executes the received command and returns the result to the user.

[1123] Step 9:

[1124] The terminal displays the execution results of commands received from the server to the user. The user can then obtain the necessary information.

[1125] Through the above processing steps, users can efficiently perform a series of operations, from natural language input to the generation and execution of appropriate scripts and commands. This system makes it easy to generate and execute complex commands, even for those lacking technical knowledge.

[1126] (Example 1)

[1127] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1128] In conventional systems, users needed knowledge of commands and scripts to perform the desired operations, making them difficult for non-specialized users. Furthermore, generating scripts and commands had to be done manually, which was time-consuming and laborious. Therefore, there was a need for a system that would allow users to easily generate commands using natural language.

[1129] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1130] In this invention, the server includes an input means for the user to input in natural language, a transmission means for the terminal to send the user's input to the server, an analysis means for the server to analyze the user's input, a generation means to generate a script or command based on the analyzed content, and a presentation means to return the script or command generated from the server to the terminal and present it to the user. This makes it possible for the user to easily generate the desired script or command in natural language and to execute operations quickly and accurately.

[1131] "Natural language" refers to the language that humans use on a daily basis, and is not a specific programming language or specialized terminology, but rather a general language.

[1132] "Input method" refers to a device or interface that allows a user to input information or instructions in natural language.

[1133] "Transmission means" refers to the communication functions and protocols that a terminal uses to send user input to a server.

[1134] "Analysis means" refers to algorithms and engines that semantically understand the natural language input received by the server and perform appropriate processing.

[1135] A "natural language processing engine" refers to software or a system that analyzes input natural language and understands its meaning and intent.

[1136] "Generation means" refers to algorithms or engines used to generate specific scripts or commands based on the analyzed content.

[1137] A "template database" refers to a set of predefined templates that a generation tool references when generating scripts or commands.

[1138] "Presentation means" refers to devices or interfaces used to display generated scripts or commands to the user.

[1139] A "terminal" refers to a computer or mobile device used by a user to input information and receive generated scripts or commands.

[1140] A "server" refers to a computer system equipped with analysis and generation capabilities, designed to process input from terminals.

[1141] Overall Overview

[1142] The present invention provides a system that automatically generates and presents necessary scripts and commands simply by having the user input instructions in natural language. This allows users to perform operations efficiently even without specialized knowledge. The system provides support for users to carry out their tasks quickly and accurately.

[1143] System Configuration

[1144] This system includes the following main components:

[1145] 1. Input method: The user inputs using natural language.

[1146] 2. Transmission method: The terminal sends the user's input to the server.

[1147] 3. Analysis method: The server analyzes the user's input.

[1148] 4. Generation method: Generate a script or command based on the analyzed content.

[1149] 5. Presentation method: The generated script or command is presented to the user.

[1150] Details of the example

[1151] 1. User input

[1152] The user inputs instructions into the terminal using natural language. For example, they might input, "Please generate a command to connect to the specified server and check disk usage." An example of a specific prompt is as follows:

[1153] Example of a prompt:

[1154] "Please generate a command to connect to the specified server and check disk usage."

[1155] "Generate a command to retrieve a list of files in a specified directory."

[1156] 2. Natural language analysis

[1157] The terminal sends the user's input to the server. The server analyzes the input using a natural language processing engine (e.g., "GPT-4"). The natural language processing engine syntactically analyzes the input text, understands its intent, and then extracts the necessary information.

[1158] 3. Script / command generation

[1159] Based on the analyzed data, the server generates appropriate scripts and commands. In this process, the server uses a command generation engine (e.g., a "shell script generation library") to generate specific commands by supplementing templates obtained from a template database. For example, to check disk usage, the Unix command "df -h" is generated.

[1160] 4. Presentation of results

[1161] The generated scripts and commands are sent back from the server to the terminal. The terminal visually presents them to the user, who can then review and use them. For example, the generated command "df -h" is displayed on the terminal screen and executed by the user.

[1162] Specific example

[1163] Example 1: Checking disk usage

[1164] User input: "Generate a command to check the server's disk usage."

[1165] Server analysis and generation: The server analyzes the input and generates the command "df -h".

[1166] Instructions: The terminal will display "df -h" to the user, who will use this to check disk usage.

[1167] Example 2: Retrieving a file list

[1168] User input: "Generate a command to get a list of files in the specified directory."

[1169] Server analysis and generation: The server generates the command "ls -la / path / to / directory".

[1170] Instructions: The terminal displays "ls -la / path / to / directory" to the user, who then uses this to obtain a list of files.

[1171] Operation flow and effects

[1172] This system allows users to quickly and easily generate appropriate scripts and commands. Even without technical knowledge, complex commands can be generated simply by inputting instructions in natural language. Therefore, users can perform their tasks efficiently, and because they can visually confirm the generated scripts and commands, they can proceed with confidence.

[1173] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1174] Step 1:

[1175] The user inputs instructions into the terminal using natural language. Specifically, the user inputs text data such as, "Please generate a command to connect to the specified server and check disk usage." This input becomes the input data required to proceed to the next step.

[1176] Step 2:

[1177] The terminal sends the user's input to the server. The terminal sends the entered text as a string to the server, which is then processed by the parsing engine. In this step, the input text is transferred from the terminal to the server.

[1178] Step 3:

[1179] The server analyzes the user's input. The server uses a natural language processing engine to analyze the input and understand the user's intent. In this process, the input data (natural language text) is syntactically analyzed and semantically. The analysis engine extracts the intent "check disk usage" from the input text and generates analysis results to proceed to the next step.

[1180] Step 4:

[1181] The server generates a script or command based on the analyzed data. Here, the server uses a command generation engine to retrieve the appropriate template from the template database and generate a specific command. For example, if the analysis result is "check disk usage", the command "df -h" will be generated. Based on the analysis result (parsed data), the generation engine processes the data and generates a specific command as output.

[1182] Step 5:

[1183] The server sends the generated scripts and commands back to the terminal. Specifically, the server sends the command "df -h" generated by the generation engine to the terminal, and a process of receiving it takes place.

[1184] Step 6:

[1185] The terminal presents the generated script or command to the user. The terminal displays the received command in the user interface, allowing the user to review and use it. For example, the command "df -h" might appear on the screen, and the user can copy and execute that command.

[1186] (Application Example 1)

[1187] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1188] Traditionally, supporting the operation of factory robots often requires advanced technical knowledge and specialized skills to properly execute complex operations and maintenance instructions. This presents challenges, particularly in responding quickly and accurately to emergencies and non-routine tasks. It also increases the risk of operational errors and work delays.

[1189] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1190] In this invention, the server includes an input means for the user to input information in natural language, an analysis means for analyzing the user's input, a generation means for generating a script or command based on the analyzed content, a presentation means for presenting the generated script or command to the user, and an execution means for applying the generated script or command to a machine. This enables the user to quickly and accurately operate and maintain factory robots without requiring specialized knowledge.

[1191] "Input means" refers to functions and devices that allow users to input commands and instructions in natural language.

[1192] "Analysis means" refers to functions and devices for analyzing the content of natural language obtained from input means and understanding the user's intent.

[1193] "Generation means" refers to a function and device for generating an appropriate script or command based on the content analyzed by the analysis means.

[1194] "Presentation means" refers to functions and devices for presenting scripts or commands generated by the generation means to the user visually or by other means.

[1195] "Execution means" refers to functions and devices for actually applying the script or command presented by the presentation means to a machine and performing operations or tasks.

[1196] A "generative AI model" is an artificial intelligence model that analyzes natural language and generates appropriate scripts or commands based on the analysis results.

[1197] This invention provides a system that streamlines the operation of factory robots, enabling users to generate and execute appropriate operation scripts and commands simply by inputting instructions in natural language. This system includes the following main components:

[1198] An "input method" refers to a function or device that allows a user to input commands or instructions in natural language. Specifically, this includes devices such as smartphones, tablets, and factory terminals. The user inputs instructions to the device in natural language, such as "Return the robot arm to the home position."

[1199] "Analysis means" refers to functions and devices for analyzing the content of natural language obtained from input means and understanding the user's intent. This system uses a generative AI model (e.g., OpenAI GPT-3) as the analysis engine. This model analyzes the input natural language using advanced algorithms and extracts the intended operation.

[1200] The "generation means" refers to a function and device for generating an appropriate script or command based on the content analyzed by the analysis means. If the analyzed content is an instruction to "return the robot arm to the home position," the generation AI model will generate a script such as "moveToHomePosition();". The generation means is mainly executed on a server, and the generated script is processed within the server.

[1201] "Presentation means" refers to functions and devices for presenting the script or command generated by the generation means to the user visually or by other means. Specifically, it displays the generated script on the screen of a smartphone or tablet. The user can then review its contents and decide whether or not to execute it.

[1202] "Execution means" refers to the functions and devices that actually apply the script or command presented by the presentation means to the machine and perform the operation or task. Once the user confirms the displayed script and approves its execution, the script is sent to the factory robot and the actual operation begins.

[1203] Specific example

[1204] For example, if a user enters "Get the current position of the robot arm" into a tablet, the system's analysis mechanism analyzes the input and uses a generation AI model to generate the command "getCurrentPosition();". The generated command is then displayed on the tablet screen for the user to confirm. After the user confirms, the command is sent to the robot, which obtains its current position and returns that information.

[1205] Example of a prompt

[1206] The following are examples of specific prompt statements used for generative AI models.

[1207] Based on the following natural language instructions, generate a script to operate a factory robot:

[1208] Instructions: Return the robot arm to its home position.

[1209] script:

[1210] In this way, by allowing users to input instructions in natural language and automatically generating appropriate scripts and commands based on those instructions, the operation of factory robots is simplified, and efficient work is achieved.

[1211] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1212] Step 1:

[1213] The user inputs a natural language instruction, such as "Return the robot arm to the home position," into an input device such as a tablet or smartphone. This input information is then transmitted to the system's input method.

[1214] Input: User's natural language instructions

[1215] Output: Instruction data to the input means

[1216] Step 2:

[1217] Natural language instructions received through the input device are sent to the server. The server's analysis device inputs these instructions into a generative AI model (e.g., OpenAI GPT-3) and analyzes the content. The generative AI model analyzes the input natural language and processes the data to understand the user's intent.

[1218] Input: Natural language instruction data

[1219] Output: Analyzed instruction data

[1220] Step 3:

[1221] The generating AI model generates an appropriate script or command based on the analyzed instruction data. For example, in response to the instruction "Return the robot arm to the home position," the script "moveToHomePosition();" is generated. This generated script is then processed by the generation mechanism.

[1222] Input: Analyzed instruction data

[1223] Output: Generated script or command

[1224] Step 4:

[1225] The generated script or command is visually displayed to the user through a presentation mechanism. The user reviews the script on their tablet or smartphone screen. At this point, the user reviews the content of the generated command and decides whether or not to execute it.

[1226] Input: Generated script or command

[1227] Output: Displayed script

[1228] Step 5:

[1229] After the user reviews the presented script and presses the execute button, the script is sent to the factory robot via the server. This causes the robot to begin its actual operation via the execution mechanism. For example, the robot arm might perform an action to return to its home position.

[1230] Input: User verification and execution command

[1231] Output: Actual operation by the robot

[1232] In this process, users can quickly and accurately operate factory robots simply by inputting instructions in natural language. By coordinating the generative AI model, presentation method, and execution method, it is possible to significantly reduce the burden on the user and achieve efficient work.

[1233] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1234] Overall Overview

[1235] The present invention provides a system that automatically generates and presents necessary scripts and commands simply by having the user input instructions in natural language. Furthermore, the present invention incorporates an emotion engine that recognizes the user's emotions, enabling it to provide appropriate responses according to the user's feelings. This system provides support for users to carry out their tasks quickly and accurately.

[1236] System Configuration

[1237] This system includes the following main components:

[1238] 1. Input method: The user inputs using natural language.

[1239] 2. Analysis method: Analyzes the user's input.

[1240] 3. Generation method: Generate a script or command based on the analyzed content.

[1241] 4. Presentation method: The generated script or command is presented to the user.

[1242] 5. Emotion Engine: Recognizes the user's emotions and provides feedback to analysis and generation methods.

[1243] Details of the example

[1244] 1. User Input and Sentiment Recognition

[1245] The user inputs instructions to the terminal using natural language. For example, if the user inputs, "I want to generate a command to connect to a specified server and check disk usage," the emotion engine simultaneously recognizes the user's emotions (e.g., frustration or tension) from the user's input, voice, facial expressions, and gestures.

[1246] 2. Analysis of Natural Language and Sentiment

[1247] The terminal sends user input and emotional information to the server. The server uses a natural language processing engine to analyze the input and an emotional processing engine to analyze the user's emotions. The analysis results include the user's desired objective (e.g., "check disk usage") and the user's emotional state.

[1248] 3. Script / command generation and emotion reflection

[1249] Based on the analyzed data, the server generates appropriate scripts and commands. For example, it generates the Unix command "df -h" to check disk usage. Simultaneously, the emotion engine generates additional information and warning messages to alleviate user frustration and tension (e.g., "Executing this command will show you your disk usage").

[1250] 4. Presentation of results and implementation

[1251] The server generates scripts, commands, and emotionally sensitive messages, which are then sent back to the terminal. The terminal presents these to the user. The user can then review the generated commands and additional information and use them.

[1252] Specific example

[1253] Example 1: Checking disk usage

[1254] User input: The user instructs the terminal to "generate a command to check the server's disk usage."

[1255] Server analysis and generation: The server analyzes the user's instructions and generates the disk usage command "df -h". The emotion engine detects the user's tension and generates an additional message: "Executing this command will show you your disk usage."

[1256] Prompt: The terminal prompts the user with the command "df -h" and an additional message. The user uses the provided command to check the server's disk usage.

[1257] Example 2: Retrieving a file list

[1258] User input: The user instructs the terminal to "generate a command to retrieve a list of files in the specified directory."

[1259] Server analysis and generation: The server analyzes the data, such as "ls -la / path / to / directory", and generates the appropriate command. The sentiment engine detects user frustration and generates an additional message such as, "Executing this command will show you a list of files in the specified directory."

[1260] Presentation: The terminal presents the user with the additional message "ls -la / path / to / directory", and the user executes the command to obtain a list of files.

[1261] Operation flow and effects

[1262] This system allows users to quickly and easily generate appropriate scripts and commands, and receive responses that take their emotions into consideration thanks to its emotion engine. Users not only input instructions in natural language, but the system also enhances the user experience by providing additional information and warning messages based on the user's emotions. Even those lacking technical knowledge can use the system without difficulty, allowing for confident operation.

[1263] The following describes the processing flow.

[1264] Step 1:

[1265] The user enters instructions into the terminal using natural language. For example, they might enter, "I want to generate a command to connect to a specified server and check disk usage."

[1266] Step 2:

[1267] The emotion engine recognizes the user's emotions from their input, voice, facial expressions, and other factors. For example, it can detect user frustration based on the tone of their voice and input speed.

[1268] Step 3:

[1269] The device sends user instructions in natural language and emotional information to the server. The data sent to the server consists of natural language text data and emotional information generated by an emotion engine.

[1270] Step 4:

[1271] The server uses a natural language processing engine to analyze the user's instructions. The processing engine identifies the user's desired objective (for example, "check disk usage").

[1272] Step 5:

[1273] The emotion engine analyzes the transmitted emotion information to identify the user's emotional state. For example, it determines whether the user is feeling "stressed" or "frustrated."

[1274] Step 6:

[1275] The server generates an appropriate script or command based on the analyzed information. For example, it might generate the Unix command "df -h" to check disk usage.

[1276] Step 7:

[1277] The server generates additional messages tailored to the user's emotions based on the analysis results of the emotion engine. For example, if the user is feeling frustrated, it will generate encouraging and reassuring messages such as, "Running this command will show you your disk usage."

[1278] Step 8:

[1279] The server sends the generated script or command and any additional messages to the terminal.

[1280] Step 9:

[1281] The terminal displays the received script or command and any additional messages to the user. The user reviews the displayed content on the terminal screen.

[1282] Step 10:

[1283] If a user needs to execute a presented script or command, they send the command to the server via the terminal. The server executes the received command and returns the result to the user.

[1284] Step 11:

[1285] The terminal displays the execution results of commands received from the server to the user. The user can then obtain the necessary information.

[1286] Through the processing steps described above, users can efficiently perform a series of operations, from natural language input to the generation of appropriate scripts and commands, to emotionally sensitive responses and execution. This system allows users to carry out their work quickly and easily, and the emotional engine provides a less stressful user experience.

[1287] (Example 2)

[1288] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1289] Current scripting and command generation systems require users to possess a high level of technical knowledge and have limited understanding of natural language instructions, thus limiting user convenience. Furthermore, they fail to provide responses that consider the user's emotional state, resulting in a lack of appropriate support, especially for users experiencing frustration or anxiety. Therefore, there is a need for a system that can quickly and accurately understand user instructions and respond with consideration for the user's emotions.

[1290] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1291] In this invention, the server includes an input means for the user to input in natural language, an analysis means for analyzing the user's input and emotional information, a generation means for generating a script or command based on the analyzed content and emotional information, and a presentation means for presenting the generated script or command and additional emotionally sensitive information to the user. As a result, the user can quickly and accurately generate the necessary scripts and commands simply by inputting instructions in natural language, and the system provides responses that are sensitive to the user's emotions, thereby improving the user experience.

[1292] "Input means" refers to a device or interface for a user to input instructions in natural language.

[1293] "Analysis means" refers to a device or software for analyzing user input and emotional information.

[1294] "Generation means" refers to a device or software for generating scripts or commands based on analyzed content and emotional information.

[1295] "Presentation means" refers to a device or interface for presenting a generated script or command and additional information that takes emotions into consideration to the user.

[1296] A "natural language processing engine" is software or an algorithm that analyzes natural language input by a user to understand its intent and purpose.

[1297] An "emotion analysis engine" is software or an algorithm that analyzes a user's emotional state from inputs such as voice, facial expressions, and gestures.

[1298] A "script" is a set of commands or program code generated to automate a specific task or operation.

[1299] A "command" is an instruction generated to tell a system to perform a specific operation or task.

[1300] "Additional information" refers to supplementary messages or explanations provided alongside generated scripts or commands, which are designed with the user's feelings in mind.

[1301] Overall Overview

[1302] The system of this invention automatically generates and presents necessary scripts and commands simply by the user inputting in natural language. Furthermore, the invention incorporates an emotion engine that recognizes the user's emotions, enabling it to provide appropriate responses according to the user's feelings. This system provides support that allows users to carry out their tasks quickly and accurately, even if they lack technical knowledge.

[1303] System Configuration

[1304] This system includes the following main components:

[1305] 1. Input means: A device or interface in which the user inputs information in natural language.

[1306] 2. Analysis means: A device or software that analyzes user input and emotional information.

[1307] 3. Generation means: A device or software that generates a script or command based on the analyzed content and emotional information.

[1308] 4. Presentation means: A device or interface that presents the generated script or command and additional information that takes emotions into consideration to the user.

[1309] Hardware and software to be used

[1310] The following hardware and software will be used for the specific implementation of the system:

[1311] Natural language processing engine: For example, software such as NLTK or OpenAI's GPT-4 can be used.

[1312] Sentiment analysis engine: For example, software such as Google's Dialogflow can be used.

[1313] Device: A computer, smartphone, or other device that the user directly operates.

[1314] Server: A remote server for executing analysis or generation methods.

[1315] Specific examples of how the system works

[1316] User input and emotion recognition

[1317] The user inputs instructions to the terminal using natural language. For example, they might input, "I want to generate a command to connect to the specified server and check disk usage." The terminal also recognizes the user's emotions (e.g., frustration or tension) from their input, voice, facial expressions, and gestures.

[1318] Analysis of natural language and emotions

[1319] The terminal sends user input and emotion information to the server. The server uses a natural language processing engine to analyze the input and an emotion engine to analyze the user's emotions. The analysis results include the user's desired objective (e.g., "check disk usage") and the user's emotional state.

[1320] Script / command generation and emotion reflection

[1321] Based on the analyzed data, the server generates appropriate scripts and commands. For example, it generates the Unix command "df -h" to check disk usage. At the same time, the emotion engine generates additional information to alleviate user frustration and tension (for example, "Executing this command will show you your disk usage").

[1322] Presentation of results and execution

[1323] The server generates scripts, commands, and emotionally sensitive messages, which are then sent back to the terminal. The terminal presents these to the user. The user can then review the generated commands and additional information and use them.

[1324] Example of a prompt

[1325] Here are some examples of specific prompt messages:

[1326] "Please generate Unix commands to connect to the specified server and check disk usage. Considering the user's anxiety, please also include an explanation of the execution results."

[1327] "Generate a Unix command to retrieve a list of files in a specified directory. Include a message explaining how this command can be helpful if the user is experiencing frustration."

[1328] In this way, the present invention realizes a system that flexibly and quickly analyzes the user's natural language input and provides an appropriate response that takes emotions into consideration.

[1329] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1330] Step 1:

[1331] The user enters information in natural language.

[1332] The user inputs instructions into the terminal using natural language. For example, they might input, "Generate a command to check the server's disk usage." This input is saved to the terminal as text data.

[1333] Step 2:

[1334] The device retrieves input content and emotional information.

[1335] The device acquires not only the user's input but also emotional information such as the user's voice, facial expressions, and gestures. As a result, the input data includes both the user's text data and emotional data.

[1336] Step 3:

[1337] The terminal sends the input data to the server.

[1338] The device sends collected text and sentiment data to the server. This data is packaged in JSON format or similar and transferred to the server over the network.

[1339] Step 4:

[1340] The server analyzes natural language.

[1341] The server analyzes the received text data using a natural language processing engine (such as NLTK or OpenAI's GPT-4). Specifically, it analyzes the user's instructions and understands the objective, which is "checking disk usage." This analysis result is then generated as intermediate data.

[1342] Step 5:

[1343] The server analyzes emotional information.

[1344] The server uses an emotion analysis engine (Google's Dialogflow) to analyze emotional data. For example, it can identify if a user is feeling nervous based on their voice tone and facial expressions. This analysis result is also generated as intermediate data.

[1345] Step 6:

[1346] The server generates commands or scripts.

[1347] The server integrates natural language processing results and sentiment analysis results to generate appropriate scripts and commands. For example, it generates the Unix command "df -h" to check disk usage. This generated command is then output.

[1348] Step 7:

[1349] The server generates additional information that takes emotions into consideration.

[1350] Based on the sentiment analysis results, the server generates additional information and explanations that take the user's emotions into consideration. For example, it might generate a message such as, "Executing this command will show you your disk usage." The command message with this additional information is then output.

[1351] Step 8:

[1352] The server sends the generated data to the terminal.

[1353] The server sends the generated scripts, commands, and additional information to the terminal. This data is also structured in JSON format or similar and transferred to the terminal over the network.

[1354] Step 9:

[1355] The device presents to the user

[1356] The terminal presents the received scripts, commands, and any additional messages to the user. The user reviews them and, if necessary, executes the commands directly on the terminal.

[1357] (Application Example 2)

[1358] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1359] Conventional factory robot systems require staff to perform complex coding and command input when issuing operating instructions to the robots, thus demanding specialized knowledge. Furthermore, a lack of consideration for staff emotional states often leads to stress and frustration. There is a need to solve these problems and make factory robot operation more efficient and user-friendly.

[1360] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an input means for the user to input in natural language, an analysis means for analyzing the user's input content and emotional state, a generation means for generating a script or command based on the analyzed content and emotional state, and a presentation means for presenting the generated script or command and an emotionally sensitive response message to the user. As a result, factory staff can easily instruct robot operation in natural language, and emotionally sensitive responses are provided, thereby reducing staff stress and enabling efficient and smooth operation.

[1361] "Input means" refers to a device or interface for a user to input instructions in natural language.

[1362] "Analysis means" refers to a device or program for analyzing the user's input and emotional state.

[1363] "Generation means" refers to a device or program for generating a script or command based on the analyzed content and emotional state.

[1364] "Presentation means" refers to a device or interface for presenting a generated script or command and an emotionally sensitive response message to the user.

[1365] An "emotion recognition engine" is a program or device that analyzes a user's facial expressions and voice tone to identify their emotional state.

[1366] A "natural language processing engine" is a program or device used to analyze the content of natural language input by a user.

[1367] A "server" is a central processing unit that executes analysis and generation methods and provides the results to the user.

[1368] A "script" is a series of instructions generated to automate a specific task.

[1369] A "command" is an instruction issued to perform a specific operation or task.

[1370] A "response message" is additional information or a warning message provided with consideration for the user's emotional state.

[1371] This invention is a system in which a user inputs instructions in natural language, generates a script or command based on the analyzed content, and then presents the generated result. The invention incorporates an emotion recognition engine that recognizes the user's emotions, thereby enabling it to provide appropriate responses according to the user's emotions.

[1372] System Overview

[1373] This system includes the following main components:

[1374] 1. Input method: The user inputs information using natural language. This includes smartphone apps and factory robot interfaces.

[1375] 2. Analysis Method: The user's input and emotional state are analyzed. This includes the Google Cloud Natural Language API and the emotion recognition engine.

[1376] 3. Generation method: Generates a script or command based on the analyzed content and emotional state. OpenAI's generative AI model is an example of this.

[1377] 4. Presentation method: The generated script or command and an emotionally sensitive response message are presented to the user. This includes smartphone apps and robot display devices.

[1378] Program processing

[1379] The server first receives natural language instructions entered by the user. Next, it uses an emotion recognition engine to analyze the user's facial expressions and tone of voice to identify their emotional state. The analysis method uses the Google Cloud Natural Language API to analyze and understand the user's instructions.

[1380] In the generation process, OpenAI's generation AI model generates appropriate scripts or commands based on the analysis results and emotional state. Simultaneously, response messages corresponding to the user's emotional state are also generated. For example, if the user is feeling stressed, a reassuring message will be generated.

[1381] The generated script or command and response messages are presented to the user via a smartphone app or the robot's display device. The user can then proceed with their task using the presented script or command.

[1382] Specific examples and prompt statements

[1383] For example, consider a scenario where a user instructs a factory line to stop.

[1384] User instruction: "Stop factory line 1."

[1385] Example of a prompt:

[1386] Effect: Stop factory line 1

[1387] Emotion Score: -0.5

[1388] Generate the appropriate command and supportive message.

[1389] In this case, the server generates a response message along with the command "stop line 1" that reads, "The operation is complete. Please let us know if there is anything else we can help you with."

[1390] This system allows users to intuitively control robots using natural language without having to input complex commands, enabling them to work efficiently and without experiencing stress during the process.

[1391] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1392] Step 1:

[1393] The user inputs instructions. The user inputs instructions in natural language using a smartphone app or the factory robot's interface. The entered natural language data is sent to the server.

[1394] Step 2:

[1395] The server receives the input. The server receives natural language data sent by the user and stores its contents. The input is natural language text data. The output is the storage and management of this text data.

[1396] Step 3:

[1397] The system performs emotion analysis using an emotion recognition engine. The server passes the input natural language data to the Google Cloud Natural Language API, which analyzes the user's emotional state from voice tone and facial expression data (which may be obtained via webcam or microphone). The input consists of natural language text data and voice / facial expression data, and the output is an emotional state score.

[1398] Step 4:

[1399] This system performs natural language processing. The server passes input data to the Google Cloud Natural Language API for language analysis. The input is natural language text data, and the output is the analyzed content (text structure and semantic information as a result of the analysis).

[1400] Step 5:

[1401] The server generates scripts and response messages. It passes the analysis results and emotional state scores to OpenAI's generative AI model, which then generates appropriate scripts or commands and emotionally sensitive response messages. The input is the analysis results and emotional state scores, and the output is the generated scripts or commands and response messages.

[1402] Step 6:

[1403] The generated results are presented. The server sends the generated script or command and response message to the user's terminal. The user receives the presented results through a smartphone app or a robot's display device. The input is the generated script or command and response message, and the output is the display on the user interface.

[1404] Step 7:

[1405] User verification and execution. The user reviews the presented script or command and response messages and uses them to proceed with the task. The user provides further input as needed, and the process continues.

[1406] These steps enable users to easily generate and execute scripts or commands based on emotionally sensitive natural language instructions.

[1407] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1408] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1409] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1410] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1411] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1412] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1413] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1414] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1415] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1416] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1417] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1418] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1419] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1420] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1421] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1422] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1423] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1424] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1425] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1426] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1427] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1428] The following is further disclosed regarding the embodiments described above.

[1429] (Claim 1)

[1430] An input method in which the user inputs in natural language,

[1431] An analysis means for analyzing user input,

[1432] A generation means that generates a script or command based on the analyzed content,

[1433] A means of presenting the generated script or command to the user,

[1434] A system that includes this.

[1435] (Claim 2)

[1436] The system according to claim 1, wherein the analysis means uses a natural language processing engine.

[1437] (Claim 3)

[1438] The system according to claim 1, wherein the generation means transmits the generated script or command to a server.

[1439] "Example 1"

[1440] (Claim 1)

[1441] An input method in which the user inputs in natural language,

[1442] A transmission means by which the terminal sends user input to the server,

[1443] The server includes an analysis means for analyzing user input,

[1444] A generation means that generates a script or command based on the analyzed content,

[1445] A presentation means that sends a script or command generated from the server back to the terminal and presents it to the user,

[1446] A system that includes this.

[1447] (Claim 2)

[1448] The system according to claim 1, wherein the analysis means uses a natural language processing engine.

[1449] (Claim 3)

[1450] The system according to claim 1, wherein the generation means generates the generated script or command based on a template obtained from a template database.

[1451] "Application Example 1"

[1452] (Claim 1)

[1453] An input method in which the user inputs in natural language,

[1454] An analysis means for analyzing user input,

[1455] A generation means that generates a script or command based on the analyzed content,

[1456] A means of presenting the generated script or command to the user,

[1457] An execution means for applying the generated script or command to a machine,

[1458] A system that includes this.

[1459] (Claim 2)

[1460] The system according to claim 1, wherein the analysis means uses a generative AI model.

[1461] (Claim 3)

[1462] The system according to claim 1, wherein the generation means transmits the generated script or command to a machine.

[1463] "Example 2 of combining an emotion engine"

[1464] (Claim 1)

[1465] An input method in which the user inputs in natural language,

[1466] An analysis means for analyzing user input and emotional information,

[1467] A generation means that generates a script or command based on the analyzed content and emotional information,

[1468] A presentation means for presenting the generated script or command and additional information that takes emotions into consideration to the user,

[1469] A system that includes this.

[1470] (Claim 2)

[1471] The system according to claim 1, wherein the analysis means uses a natural language processing engine and an emotion processing engine.

[1472] (Claim 3)

[1473] The system according to claim 1, wherein the generating means transmits the generated script or command and additional information that takes emotions into consideration to the presenting means.

[1474] "Application example 2 when combining with an emotional engine"

[1475] (Claim 1)

[1476] An input method in which the user inputs in natural language,

[1477] An analysis means for analyzing the user's input and emotional state,

[1478] A generation means that generates a script or command based on the analyzed content and emotional state,

[1479] A presentation means for presenting the generated script or command and an emotionally sensitive response message to the user,

[1480] A system that includes this.

[1481] (Claim 2)

[1482] The system according to claim 1, wherein the analysis means uses an emotion recognition engine and a natural language processing engine.

[1483] (Claim 3)

[1484] The system according to claim 1, wherein the generation means sends the generated script or command and a sentiment-sensitive response message to the server. [Explanation of Symbols]

[1485] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. An input method in which the user inputs in natural language, An analysis means for analyzing user input, A generation means that generates a script or command based on the analyzed content, A means of presenting the generated script or command to the user, A system that includes this.

2. The system according to claim 1, wherein the analysis means uses a natural language processing engine.

3. The system according to claim 1, wherein the generation means transmits the generated script or command to a server.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A