system
The system addresses the non-intuitiveness of RPA tools by allowing natural language input, automatic code generation, and flowchart editing, enabling non-programmers to automate tasks effectively.
Patent Information
- Application Number
- JP2024140362
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-03-06
AI Technical Summary
Conventional RPA tools require programming knowledge and are non-intuitive, making them difficult for non-programmers to use and edit.
A system that allows users to input requests in natural language, automatically generates code, visually displays it in a flowchart format, and enables intuitive editing, using a natural language processing engine and external artificial intelligence services.
Enables non-programmers to intuitively create, edit, and automate RPA tasks, improving user experience and operational efficiency.
Smart Images

Figure 2026037337000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional RPA (Robotic Process Automation) tools often require programming knowledge, making it difficult for non-programmers to use them. Furthermore, creating and modifying code is non-intuitive, placing a significant burden on users. For this reason, there is a demand for user-friendly RPA tools that can be operated intuitively. [Means for solving the problem]
[0005] To solve this problem, the present invention provides the following means: It provides an interface for users to input requests in natural language, and includes means for generating code based on the input natural language request. It also provides a system that includes means for visually displaying the generated code in a flowchart format and making it editable. This allows even non-programmers to intuitively create and edit RPA tasks. Furthermore, by using a natural language processing engine as code generation means, it is possible to automatically generate code based on user requests, and by using external artificial intelligence services, it is possible to generate advanced code.
[0006] "User" refers to an individual or organization that utilizes the system to input requests in natural language and build and edit RPA tasks.
[0007] "Natural languages" are languages that are used daily and are expressed in normal text or speech, such as Japanese and English.
[0008] An "interface" is the means by which a user interacts with a system, including input fields and chat boxes that allow natural language input.
[0009] "Code" refers to statements that describe a computer program, and in the present invention is expressed in a programming language such as Python.
[0010] "Generating" refers to automatically creating a computer program (code) based on a user's request.
[0011] "Visual display" means showing the contents of the code and the flow of processing to the user in a graphical format, including formats such as flowcharts.
[0012] A "flowchart" is a diagram that illustrates the flow of a process or system, making it possible to visualize each step.
[0013] "Making it editable" means that the user can select each step on the flowchart and make modifications or additions.
[0014] A "natural language processing engine" is software or algorithms for parsing natural language input and converting it into an appropriate form.
[0015] "Artificial intelligence service" refers to an external computer system or platform that uses technologies such as machine learning and deep learning to perform advanced analysis and generation. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes an embodiment of the present invention.
[0038] The present invention relates to a system that provides an RPA tool that can be intuitively operated by a user, including an interface for users to input requests in natural language, a means for automatically generating code based on the requests, and a means for visually displaying the generated code and making it editable in a flowchart format.
[0039] Interface
[0040] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of RPA tasks in everyday languages such as English or Japanese.
[0041] Code generation
[0042] The terminal takes user input and sends a code generation request to the server, which receives the request and uses a natural language processing engine to generate Python code based on the user's requirements. The generated code is used to automatically execute RPA tasks.
[0043] The server uses an external artificial intelligence service as its natural language processing engine. This AI service analyzes natural language requests entered by users and converts them into appropriate commands and operational procedures in code. The generated code can then perform a variety of RPA tasks, such as accessing and logging into websites, retrieving data, and saving screenshots.
[0044] Code visualization and editing
[0045] The terminal visually displays the code received from the server, specifically showing the generated Python code to the user in the form of a flowchart, which illustrates how each step is processed and allows the user to intuitively understand it.
[0046] Users can check the code based on the flowchart and make any necessary modifications by clicking on a specific step in the flowchart and entering the details of the modification. For example, a user can make a modification such as "add a step to click on a specific menu after logging in."
[0047] The device then sends the user's modifications back to the server. The server then regenerates the code based on the new modifications and returns the new modified code to the device. The flowchart is updated again on the device, allowing the user to visually confirm the changes.
[0048] Specific examples
[0049] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following Python code based on the user's request (the actual content of the code is not shown here):
[0050] The generated code contains instructions to open a web browser, access the specified URL, enter a username and password to log in, take a screenshot of the dashboard page, and save it to a local file. This code is then sent back to the device and displayed as a flowchart.
[0051] The user can review the flowchart and request modifications, such as adding a step to click a specific menu after logging in. The terminal sends the modified request to the server, which then generates new code and returns it to the terminal. As a result, the RPA task required by the user can be fully executed.
[0052] The above is one embodiment of the present invention, which enables even non-programmers to intuitively create, edit, and automate RPA tasks.
[0053] The processing flow will be explained below.
[0054] Step 1:
[0055] The user accesses the system interface. The user logs in to the system using a web browser or a dedicated application.
[0056] Step 2:
[0057] The user enters a request for an RPA task into the interface in natural language, such as "Open a specific website every morning at 9am, log in, and take a screenshot of the dashboard."
[0058] Step 3:
[0059] The terminal receives the user's request and sends a code generation request to the server, which includes the user's input.
[0060] Step 4:
[0061] The server receives the request and uses a natural language processing engine to generate Python code based on the user's requirements. The server uses an external artificial intelligence service to analyze the natural language and generate the appropriate code.
[0062] Step 5:
[0063] The server sends the generated Python code in JSON format back to the device, which receives the response and parses the code.
[0064] Step 6:
[0065] It visually displays the code received by the terminal, specifically showing the generated Python code to the user in the form of a flowchart, illustrating how each step is processed.
[0066] Step 7:
[0067] The user reviews the flowchart and modifies or adds code as needed, clicking on a specific step and entering the modifications into the interface.
[0068] Step 8:
[0069] The device receives the user's correction request and sends it back to the server, containing the current code and the corrections.
[0070] Step 9:
[0071] The server receives the code generation request again and uses a natural language processing engine to generate new code based on the modifications.
[0072] Step 10:
[0073] The server generates a new code and sends it back to the device, which receives the new code and displays it again as a flowchart.
[0074] Step 11:
[0075] The user checks the updated flowchart to ensure that the changes are reflected correctly, and makes further changes if necessary.
[0076] Step 12:
[0077] Finally, once the RPA task is constructed to the user's satisfaction, the terminal executes the generated Python code, allowing the user to review the automated task and evaluate the results.
[0078] These are the program processing steps, which allows even non-programmers to intuitively create, edit, and automate RPA tasks.
[0079] Example 1
[0080] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0081] Conventional RPA tools are rarely intuitive for non-programmers to use, and require specialized knowledge, especially for generating and modifying code. Therefore, there is a need for a system that allows for input of instructions in natural language, visual code display, and editing in flowchart format. There is also a need for a system that allows for intuitive and easy modification of generated code.
[0082] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0083] In this invention, the server includes means for providing an interface for a user to input requests in natural language, means for generating code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, means for a terminal to resend user modifications to the server, the server regenerating code based on the modifications, and means for resending the resulting generated code to the terminal and updating the flowchart. This enables even non-programmers to intuitively create, edit, and automate RPA tasks.
[0084] An "interface" is a means through which a user can input requests to a system in natural language.
[0085] A "natural language request" refers to a user entering instructions or requests in everyday language.
[0086] The "code generation means" is a means having the function of automatically generating executable code based on a natural language request input by a user.
[0087] A "natural language processing engine" is an algorithm or model that analyzes natural language text, understands its content, and performs appropriate processing.
[0088] "Artificial intelligence services" refers to a general range of services that perform advanced natural language processing and machine learning using external cloud services or specific artificial intelligence models.
[0089] A "flowchart" is a diagram that visually shows each step and logic of code, and is presented in a format that is easy for users to understand intuitively.
[0090] "Terminal" refers to a device that a user uses to access and operate the system, including PCs, smartphones, tablets, etc.
[0091] A "server" is a central computer system that receives requests from users and performs code generation and other processing.
[0092] "Modifications" refers to any changes or additional instructions you make to the generated code.
[0093] A "code generation request" is a communication request to ask the server to generate code based on a user's natural language request.
[0094] This invention relates to a system that generates, modifies, and visualizes RPA task code based on user input in natural language. Specific hardware and software usage methods are described below.
[0095] Interface
[0096] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. For example, this could be a web browser such as GOOGLE CHROME® or Mozilla Firefox, or a specific application. Users can enter details of RPA tasks in everyday language such as English or Japanese.
[0097] Code generation
[0098] The terminal receives user input and sends a code generation request to the server. For example, a web application written in Python is executed via the terminal. The server receives this request and generates code based on the user's request using a natural language processing engine. The natural language processing engine includes an external generative AI model (e.g., GPT-3 (registered trademark) from OpenAI (registered trademark)). The generated code is used to automatically perform RPA tasks, such as accessing and logging in to a website, retrieving data, and saving screenshots.
[0099] Code visualization and editing
[0100] The terminal visually displays the code received from the server. Specifically, it shows the generated Python code to the user in the form of a flowchart. The flowchart illustrates how each step is processed, allowing the user to intuitively understand it. The user can review the code based on this flowchart and make corrections as necessary. For example, they can click on a specific step in the flowchart and enter corrections. Corrections could include adding a step to click a specific menu after logging in.
[0101] Resend your fixes and update your code
[0102] The device then sends the user's modifications back to the server. The server then generates the code again based on the new modifications and returns the new modified code to the device. The flowchart is updated again on the device, allowing the user to visually confirm the changes.
[0103] Specific examples
[0104] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following Python code based on the user's request (the actual content of the code is omitted here, but the generated content is sent from the server to the terminal as is). The generated code includes steps to open a web browser, access the specified URL, enter the username and password to log in, and take a screenshot of the dashboard page and save it to a local file.
[0105] Prompt Sentence Examples
[0106] Here are some examples of prompts to input to a generative AI model:
[0107] "Every morning at 9am, open a specific website, enter your username and password, log in, take a screenshot of the dashboard, and save it to a local file."
[0108] This system allows even non-programmers to intuitively create, edit, and automate RPA tasks.
[0109] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0110] Step 1:
[0111] Users access the system's interface using a web browser or a dedicated application.
[0112] Input: User launches a web browser or dedicated application
[0113] Output: The system interface is displayed.
[0114] Specific operation: The user launches a browser on their PC or mobile device and accesses the system's URL. In the case of a dedicated application, the user connects to the interface by launching the app.
[0115] Step 2:
[0116] The user enters the details of the RPA task in natural language into input fields displayed in the interface.
[0117] Input: User inputs request in natural language
[0118] Output: A natural language request is sent to the terminal
[0119] Specific actions: The user enters information such as "Open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard," and clicks the submit button.
[0120] Step 3:
[0121] The terminal obtains the user's input and sends a code generation request to the server.
[0122] Input: A natural language request entered by the user.
[0123] Output: Code generation request to the server
[0124] Specific operation: The terminal converts the information entered by the user into a data format such as JSON, and sends the converted data to the server as an HTTP request.
[0125] Step 4:
[0126] The server performs natural language processing based on the received request and generates Python code.
[0127] Input: Code generation request sent from the device
[0128] Output: Generated Python code
[0129] How it works: The server sends a request to an external generative AI model (e.g., GPT-3), which analyzes the user's request and generates a code. The generated code is sent back to the server, which then organizes it and sends it to the device.
[0130] Step 5:
[0131] The terminal visually displays the Python code received from the server in a flowchart format.
[0132] Input: Python code sent from the server
[0133] Output: Visual representation in the form of a flowchart
[0134] Specific operation: The terminal parses the received Python code and converts each step into a flowchart node. The generated flowchart is displayed to the user.
[0135] Step 6:
[0136] The user checks the flowchart and makes corrections as necessary.
[0137] Input: Visually displayed flowchart
[0138] Output: Corrected steps or instructions
[0139] Specific operation: The user clicks on a specific step on the flowchart and enters the corrections. For example, the user can make corrections such as "add a step to click on a specific menu after logging in."
[0140] Step 7:
[0141] The terminal retransmits the user's modifications to the server, and the server generates the code again based on the modifications.
[0142] Input: The corrections entered by the user
[0143] Output: Updated Python code
[0144] Specific operation: The device converts the corrections entered by the user back into JSON format or similar and sends the converted data as an HTTP request to the server. The server then sends a new request to the AI model to generate new code. The new code is then sent back to the device, which then updates the flowchart.
[0145] In this way, even non-programmers can intuitively create, edit, and automate RPA tasks.
[0146] (Application example 1)
[0147] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0148] Inventory management at distribution centers requires a high degree of efficiency and accuracy, but often relies on manual processes and legacy systems. This can lead to problems such as delays in inventory checks and replenishment instructions, inaccurate data processing, and a lack of programming skills for automation. This results in issues such as product supply delays and excess inventory, which can reduce operational efficiency. Therefore, there is a need for automation tools that can be operated intuitively even by non-programmers.
[0149] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0150] In this invention, the server includes means for providing an interface for a user to input a request in natural language, means for generating a code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, and means for automating inventory management tasks at a logistics center based on the generated code. This makes it possible for even non-programmers to easily automate inventory management tasks and manage inventory efficiently and accurately.
[0151] "User" refers to a person who utilizes the system and enters requests in natural language.
[0152] "Natural language request" refers to the specific instructions or requests entered by the user in a natural language such as English or Japanese.
[0153] "Interface" refers to the mechanisms, such as screens and input fields, provided for users to input requests in natural language.
[0154] "Code generation means" refers to a function that automatically generates programming code based on natural language requests entered by a user.
[0155] "Flowchart format" refers to a format that visually shows the generated programming code and illustrates the flow of each step.
[0156] A "logistics center" refers to a facility that stores goods, manages inventory, processes orders, etc.
[0157] "Inventory management tasks" refer to specific inventory-related tasks such as checking inventory, issuing replenishment orders, and data management.
[0158] "Automation" refers to the automatic execution of tasks that were previously performed manually using machines or software.
[0159] A "natural language processing engine" refers to technology that analyzes natural language input by a user and generates appropriate programming code based on that.
[0160] "Artificial intelligence service" refers to a system that uses AI technology provided by external cloud services, APIs, etc. to perform specific tasks.
[0161] The present invention relates to a system for providing an inventory management automation tool for a distribution center that can be intuitively operated by a user. The system includes an interface for a user to input a request in natural language, a means for automatically generating a code based on the request, and a means for visually displaying the generated code and making it editable in a flowchart format.
[0162] Interface
[0163] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of inventory management tasks at the distribution center in everyday languages such as English and Japanese.
[0164] Code generation
[0165] The server takes user input and processes code generation requests. The server receives this request and uses a natural language processing engine to generate Python code based on the user's requirements. The generated code is used to automatically perform inventory management tasks at a distribution center. For example, it can automate tasks such as checking inventory every morning at 9:00, listing missing items, and issuing replenishment instructions.
[0166] The server uses an external artificial intelligence service as its natural language processing engine. This AI service analyzes natural language requests entered by users and converts them into appropriate operational commands and code. The generated code can then perform a variety of inventory management tasks, such as logging in, retrieving data, checking inventory, and issuing replenishment instructions.
[0167] Code visualization and editing
[0168] The terminal visually displays the code received from the server, specifically showing the generated Python code to the user in the form of a flowchart, which illustrates how each step is processed and allows the user to intuitively understand it.
[0169] The user checks the code based on the flowchart and makes any necessary corrections. Corrections are made by clicking on a specific step on the flowchart and entering the details of the correction. For example, a user could make a correction such as "adding a step that automatically issues a replenishment instruction when stock runs out." The terminal then sends the corrections back to the server. The server then generates a new code based on the new corrections and returns the new, corrected code to the terminal. The flowchart is updated again on the terminal, allowing the user to visually confirm the changes.
[0170] Specific examples
[0171] As a concrete example, consider the case where a user inputs a request such as "Check inventory every morning at 9:00, list the missing items, and issue a replenishment order." The server generates the following Python code based on the user's request (the actual content of the code is not shown here):
[0172] The generated code accesses the distribution center's database to check inventory status, lists any missing items, and includes instructions for automatically issuing replenishment instructions for those items. This code is returned to the terminal and displayed as a flowchart. The user can review the flowchart and request modifications, such as adding a step to automatically place an order with a specific supplier when inventory is low. The terminal then sends the modified request to the server, which then generates a new code and returns it to the terminal. As a result, the inventory management tasks required by the user are fully automated.
[0173] Prompt Sentence Examples
[0174] "Check inventory every morning at 9am, list any items that are in short supply, and issue replenishment instructions."
[0175] The above is an embodiment of the present invention, which enables even non-programmers to intuitively create, edit, and automate inventory management tasks.
[0176] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0177] Step 1:
[0178] The user inputs requests for inventory management tasks at the distribution center into the interface in natural language.
[0179] Input: Natural Language Request "Check inventory every morning at 9:00, list any missing items, and issue replenishment instructions."
[0180] Output: Request data to the server
[0181] Specific actions: The user uses a web browser or dedicated application to enter a request into a chat box or input field in the interface.
[0182] Step 2:
[0183] The server receives a request entered in natural language and analyzes the request using a natural language processing engine.
[0184] Input: User's natural language request
[0185] Output: Code generation request based on your requirements
[0186] Specific operation: The server makes an API call to a natural language processing engine (e.g., an external artificial intelligence service) to analyze the user's request.
[0187] Step 3:
[0188] The server generates Python code based on the analysis results and creates instructions to automate the steps of inventory management tasks.
[0189] Input: Analysis results from the natural language processing engine
[0190] Output: Python code for automated inventory management tasks
[0191] Specific operation: Based on the analysis results of the natural language processing engine, the server generates Python code to automatically perform specific inventory management tasks.
[0192] Step 4:
[0193] The server sends the generated Python code to the device.
[0194] Input: Generated Python code
[0195] Output: Data sent to the terminal
[0196] What it does: The server sends the generated Python code to the terminal and sends data for visualization in the form of a flowchart.
[0197] Step 5:
[0198] The terminal visually displays the received Python code and presents it to the user in the form of a flowchart.
[0199] Input: Python code sent from the server
[0200] Output: Visual representation in the form of a flowchart
[0201] What it does: The terminal uses a flowchart library (e.g., Graphviz) to parse the generated Python code and display each step in a flowchart format.
[0202] Step 6:
[0203] The user checks the flowchart and makes corrections as necessary.
[0204] Input: Code displayed in flowchart format
[0205] Output: User modifications
[0206] Specific operation: The user clicks on a specific step in the flowchart and enters the details of the modification, such as "add a step to automatically order from a specific supplier when inventory is low."
[0207] Step 7:
[0208] The device sends the modifications back to the server, which generates a new code.
[0209] Input: User modifications
[0210] Output: The newly generated Python code
[0211] What it does: The device sends the user's modifications to the server, and the server regenerates the Python code based on the new modifications.
[0212] Step 8:
[0213] The server sends the final modified code to the terminal, and the terminal redisplays the updated flowchart.
[0214] Input: Newly generated Python code
[0215] Output: Updated visual representation in the form of a flowchart
[0216] Specific operation: The server sends the final modified code to the terminal, and the terminal displays the updated code again as a flowchart.
[0217] Step 9:
[0218] The server runs the final generated code and automates inventory management tasks.
[0219] Input: Final revised Python code
[0220] Output: The results of the automated inventory management task.
[0221] What happens: The server executes the final generated code and automates inventory management tasks (stock checks, replenishment orders, etc.).
[0222] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0223] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes an embodiment of the present invention.
[0224] This invention relates to a system that allows users to input requests in natural language and automatically generates and edits RPA tasks. Furthermore, it aims to improve the user experience by combining it with an emotion engine that recognizes the user's emotions.
[0225] Interface
[0226] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of RPA tasks using text or voice.
[0227] Code generation
[0228] The device takes the user's input and sends a code generation request to the server. The server receives this request and uses a natural language processing engine to generate Python code based on the user's requirements. The server then uses an external artificial intelligence service to analyze the natural language and generate the appropriate code.
[0229] The server uses an external artificial intelligence service as its natural language processing engine. This AI service analyzes natural language requests entered by users and converts them into appropriate commands and operational procedures in code. The generated code can then perform a variety of RPA tasks, such as accessing and logging into websites, retrieving data, and saving screenshots.
[0230] Emotion Engine
[0231] This invention also incorporates an emotion engine, which analyzes emotions from the user's voice and text input and provides the analysis results to the natural language processing engine. This enables code generation and feedback according to the user's emotional state.
[0232] For example, if a user types "This task is tiring," the emotion engine will recognize the emotion "tiring" and provide it to the natural language processing engine, which can then use this information to more efficiently generate the appropriate code to improve the user's experience.
[0233] Code visualization and editing
[0234] The terminal visually displays the code received from the server, specifically showing the generated Python code to the user in the form of a flowchart, which illustrates how each step is processed and allows the user to intuitively understand it.
[0235] Users can review the code based on the flowchart and make corrections or additions as needed. They can click on specific steps and enter their corrections into the interface. The emotion engine also analyzes the user's emotions and provides accurate feedback.
[0236] Specific examples
[0237] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following Python code based on the user's request (the actual content of the code is not shown here):
[0238] The generated code contains instructions to open a web browser, access the specified URL, enter a username and password to log in, take a screenshot of the dashboard page, and save it to a local file. This code is then sent back to the device and displayed as a flowchart.
[0239] The user can review the flowchart and request modifications, such as adding a step to click a specific menu after logging in. The terminal sends the modified request to the server, which then generates new code and returns it to the terminal. As a result, the RPA task required by the user can be fully executed.
[0240] With the built-in emotion engine, even if the user inputs an emotional request such as "Is there a way to make this operation easier?", the system will make appropriate code modifications accordingly. In this way, the system can analyze the user's emotions and improve the user experience.
[0241] The above is one embodiment of the present invention, which enables even non-programmers to intuitively create and edit RPA tasks and automate them with an approach that takes user emotions into consideration.
[0242] The processing flow will be explained below.
[0243] Step 1:
[0244] The user accesses the system interface. The user logs in to the system using a web browser or a dedicated application.
[0245] Step 2:
[0246] The user enters a request for an RPA task into the interface in natural language, such as "Open a specific website every morning at 9am, log in, and take a screenshot of the dashboard."
[0247] Step 3:
[0248] The device receives the user's request and sends the data to the emotion engine, which analyzes the user's input text or voice to determine the emotion and sends the results back to the device.
[0249] Step 4:
[0250] Based on the emotion analysis results, the device sends a code generation request to the server. The request includes the user's input and the emotion analysis results.
[0251] Step 5:
[0252] The server receives the request and uses a natural language processing engine to generate Python code based on the user's request and sentiment analysis. The server uses an external artificial intelligence service to analyze the natural language and generate the appropriate code.
[0253] Step 6:
[0254] The server sends the generated Python code in JSON format back to the device, which receives the response and parses the code.
[0255] Step 7:
[0256] It visually displays the code received by the terminal, specifically showing the generated Python code to the user in the form of a flowchart, illustrating how each step is processed.
[0257] Step 8:
[0258] The user reviews the flowchart and makes any necessary code modifications or additions. They click on specific steps and enter their modifications into the interface. The emotion engine also analyzes the user's emotions and provides accurate feedback.
[0259] Step 9:
[0260] The device receives the user's correction request and sends it back to the server. The request includes the current code, the correction content, and the sentiment analysis results.
[0261] Step 10:
[0262] The server receives the code generation request again and uses a natural language processing engine to generate new code based on the modifications.
[0263] Step 11:
[0264] The server generates a new code and sends it back to the device, which receives the new code and displays it again as a flowchart.
[0265] Step 12:
[0266] The user checks the updated flowchart to ensure that the changes are reflected correctly, and makes further changes if necessary.
[0267] Step 13:
[0268] Finally, once the RPA task is constructed to the user's satisfaction, the terminal executes the generated Python code, allowing the user to review the automated task and evaluate the results.
[0269] These are the specific processing steps of a system incorporating an emotion engine. This enables even non-programmers to intuitively create and edit RPA tasks and automate them with an approach that takes user emotions into account.
[0270] Example 2
[0271] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0272] Conventional RPA (Robotic Process Automation) systems are often difficult for users without programming knowledge to operate intuitively. Furthermore, they do not provide an interface that takes user emotions into consideration, which does not lead to an improvement in the actual user experience. This often leaves users feeling dissatisfied and stressed.
[0273] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0274] In this invention, the server includes means for providing an interface for a user to input a request in natural language, means for generating code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, an emotion engine for analyzing the user's emotions, and means for adjusting the code based on the emotion analysis results. This allows even non-programmers to operate intuitively, and further provides feedback that takes the user's emotions into consideration, thereby improving the user experience.
[0275] An "interface" is the means by which a user accesses a system and inputs requests.
[0276] A "natural language request" is a request that a user enters into a system using natural language (e.g., everyday conversation or instructions).
[0277] The "code generation means" is a function for creating program code based on a natural language request input by a user.
[0278] "Visual display" means presenting the generated code in the form of diagrams or flowcharts so that the user can intuitively understand it.
[0279] "Editable in flowchart format" means that the user can visually check the generated code and make corrections or additions as needed.
[0280] The "emotion engine" is a function that analyzes emotions from user input (text or voice) and provides that emotional data to the system.
[0281] "Emotion analysis results" refers to the emotional information analyzed by the emotion engine from the user's input.
[0282] A "natural language processing engine" is a technology for analyzing a user's natural language requests and generating corresponding program code.
[0283] "External artificial intelligence services" means using AI technology or services provided by external providers rather than within the system.
[0284] The following describes an embodiment of the present invention. This invention relates to a system that allows users to input requests in natural language and automatically generates and edits RPA tasks. Furthermore, the invention aims to improve the user experience by combining it with an emotion engine that recognizes the user's emotions.
[0285] Interface
[0286] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of RPA tasks using text or voice.
[0287] Code generation
[0288] The device receives the user's input and sends a code generation request to the server. The server receives the request and uses a natural language processing engine to generate program code based on the user's request. The server then uses an external artificial intelligence service (e.g., GPT-4 (registered trademark)) to analyze the natural language and generate the appropriate code.
[0289] Emotion Engine
[0290] The present invention further incorporates an emotion engine. The emotion engine analyzes emotions from the user's voice or text input and provides the analysis results to the natural language processing engine. This enables code generation and feedback according to the user's emotional state. For example, if the user inputs "This task is tiring," the emotion engine recognizes the emotion "tiring" and provides it to the natural language processing engine. Based on this information, the natural language processing engine can more efficiently generate appropriate code to improve the user's experience.
[0291] Code visualization and editing
[0292] The device visually displays the code received from the server. Specifically, it shows the generated program code to the user in the form of a flowchart. The flowchart illustrates how each step is processed, allowing the user to intuitively understand. The user can review the code based on the flowchart and make corrections or additions as needed. When the user clicks on a specific step and enters the corrections into the interface, the emotion engine analyzes the user's emotions and provides accurate feedback.
[0293] Specific examples
[0294] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following program code based on the user's request (details of the code are omitted here).
[0295] The generated code contains instructions to open a web browser, access the specified URL, enter a username and password to log in, take a screenshot of the dashboard page, and save it to a local file. This code is then sent back to the device and displayed as a flowchart.
[0296] Users can review the flowchart and request modifications, such as adding a step to click a specific menu after logging in. The terminal sends the modified request to the server, which then generates new code and returns it to the terminal. As a result, the RPA task required by the user can be fully executed.
[0297] With the built-in emotion engine, even if the user inputs an emotional request such as "Is there a way to make this operation easier?", the system will make appropriate code modifications accordingly. In this way, the system can analyze the user's emotions and improve the user experience.
[0298] This invention enables even non-programmers to intuitively create and edit RPA tasks and automate them with an approach that takes user emotions into consideration.
[0299] The above is a specific embodiment for carrying out the present invention.
[0300] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0301] Step 1:
[0302] The user accesses the interface
[0303] A user accesses the system's interface using a web browser or a dedicated application. The input is the launch of the web browser or dedicated application. The output is the display of the interface, waiting for user input.
[0304] Step 2:
[0305] The user enters a request
[0306] The user inputs the details of the RPA task into the interface in natural language, either through text or voice. The input is the user's natural language request, for example, a prompt such as "Open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The output is the user's request data, which is sent to the interface.
[0307] Step 3:
[0308] The device gets user input
[0309] The terminal receives user input from the interface. As input, it receives the user's natural language request data. The terminal then prepares the request data for parallel processing. As output, it formats the request data in a well-formed way and sends it to the server.
[0310] Step 4:
[0311] The device sends a code generation request to the server
[0312] The terminal sends the user's request, properly formatted, to the server. As input, the formatted request data is sent. As output, a code generation request is issued to the server.
[0313] Step 5:
[0314] The server receives the request
[0315] The server receives a request sent from the terminal. The input is the code generation request data. The server then prepares to analyze the request data. The output is a state ready for analysis.
[0316] Step 6:
[0317] The server starts the natural language processing engine
[0318] The server launches an external AI service (e.g., GPT-4) to analyze the user's request. The input is the user's natural language request data. The output is intermediate data of the analysis results.
[0319] Step 7:
[0320] The server requests analysis from the emotion engine
[0321] The server sends the user's input data to the emotion engine and requests emotion analysis. The input is the user's input data. The output is the emotion analysis result returned from the emotion engine.
[0322] Step 8:
[0323] The emotion engine analyzes user input
[0324] The emotion engine analyzes emotions from the user's text or voice input. The input is the user's input data. Through the analysis process, the user's emotional state (e.g., "fed up") is extracted as data. The output is the emotion analysis result returned to the server.
[0325] Step 9:
[0326] The server generates code based on the analysis results
[0327] The server generates program code based on the analysis results obtained from the natural language processing engine and emotion engine. The inputs are the natural language analysis results and emotion analysis results. The output is Python code that meets the user's requirements.
[0328] Step 10:
[0329] The server sends the generated code to the device.
[0330] The server sends the generated Python code to the terminal. The input is the generated Python code. The output is the code sent to the terminal.
[0331] Step 11:
[0332] The terminal will visually display the code
[0333] The terminal visually displays the code received from the server, specifically in the form of a flowchart, to the user. The input is the Python code sent by the server. The output is a visual representation of the flowchart provided on the interface.
[0334] Step 12:
[0335] User reviews code and requests corrections
[0336] The user checks the displayed flowchart and requests corrections or additions as necessary. The inputs are the visually displayed flowchart and the user's correction requests. The output is the correction request data entered into the interface.
[0337] Step 13:
[0338] The device sends a modification request to the server
[0339] The terminal sends the user's correction request data to the server. The correction request data is the input. The output is a code generation request sent to the server again.
[0340] Step 14:
[0341] The server generates the modified code
[0342] The server generates new code based on the modification request. The input is the modification request data. The output is the updated Python code.
[0343] Step 15:
[0344] The server sends a new code to the device.
[0345] The server sends the updated Python code to the terminal. As input, it takes the updated Python code. As output, it sends the new code to the terminal.
[0346] Step 16:
[0347] The terminal will visually display the new code.
[0348] The terminal again visually displays the new code received from the server. As input, the updated Python code. As output, a visual representation of the updated flowchart is provided on the interface.
[0349] (Application example 2)
[0350] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0351] Current automation systems and robot control systems require specialized knowledge and are complex to operate, making them difficult for users to use intuitively. Furthermore, they do not take into account the user's emotions or work attitudes, resulting in problems that do not improve work efficiency or the user experience. This invention aims to provide a system that allows users to perform automated tasks by issuing instructions in natural language, and to realize a more user-friendly and efficient system by taking the user's emotions into account.
[0352] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for providing an interface for a user to input a request in natural language, means for generating code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, and emotion analysis means for analyzing the user's emotions and reflecting the analysis results in code generation. This allows the user to intuitively issue instructions in natural language, and further enables the system to generate and modify optimal code taking the user's emotional state into consideration.
[0353] An "interface" is a means by which a user can input requests in natural language.
[0354] The "code generation means" is a function that generates appropriate program code based on natural language requests entered by the user.
[0355] A "natural language processing engine" is an artificial intelligence-based system for analyzing natural language requests and generating corresponding code.
[0356] A "flowchart format" is a diagram format that visually represents each step of a program and allows users to operate it intuitively.
[0357] The "emotion analysis means" is a function that analyzes emotions from the user's voice input or text input and provides the analysis results to the system.
[0358] "External Artificial Intelligence Service" means an artificial intelligence-based service hosted on an external server that is used to analyze natural language and generate code.
[0359] A "robotic task" is a series of movements that are programmed to automate a specific task or action within a factory.
[0360] An "RPA script" is a program code written to automate software operating procedures and is used for robotic process automation.
[0361] "Emotional state" refers to the psychological state a user feels in response to a particular request or task.
[0362] The following describes the embodiments of the present invention: The system provides a set of functions for users to input requests in natural language and optimize robot operations based on sentiment analysis.
[0363] System Configuration
[0364] The system consists of a user terminal, a server, and factory robots. The user terminal provides an interface for inputting requests in natural language. The server is responsible for analyzing the natural language requests and generating the corresponding RPA script. It also analyzes the user's emotions and reflects them in the work content. Finally, the generated script is distributed to the factory robots, which then perform the specified work.
[0365] Hardware and software used
[0366] The hardware used includes users' PCs and tablets, servers (with high-performance processors and large amounts of memory), and robots that perform work in the factory.The software used includes a script generator developed in Python, an NLP (natural language processing) engine, a sentiment analysis engine, and external artificial intelligence service APIs.
[0367] Natural Language Processing and Sentiment Analysis
[0368] When a user inputs a request in natural language through their device, the request is sent to the server. The server then calls an external artificial intelligence service to analyze the input natural language. This service analyzes the user's request and generates a corresponding code. It also uses a sentiment analysis engine to analyze the user's emotions and adjust the generated code based on the results.
[0369] Visual Display and Editing
[0370] The generated code is visually displayed in a flowchart format for easy user understanding. The user can review the flowchart and edit the work steps as needed, entering modifications using the user's device interface. These modifications are then sent back to the server, where a new script is generated.
[0371] Specific examples
[0372] As a concrete example, consider the case where a user inputs "Start the line belt at 8:00 every morning and collect data from the temperature sensor." Based on this request, the system performs the following processing:
[0373] 1. A natural language processing engine analyzes the user's request and generates an RPA script to start the line belt and collect data from the temperature sensor.
[0374] 2. The sentiment analysis engine analyzes the user's emotional state and adjusts the script as needed.
[0375] 3. The generated script is displayed in flowchart format for the user to check and modify.
[0376] 4. Any modifications are processed again on the server and the final script is generated.
[0377] 5. The final script is delivered to the robot, which then performs the specified task.
[0378] This allows users to intuitively give instructions and automate tasks efficiently.
[0379] Prompt Sentence Examples
[0380] For example, if a user types "Start Linebelt every morning at 8am and collect temperature sensor data", the information passed to the system is:
[0381] User entered text: "Start Linebelt every morning at 8am and collect temperature sensor data"
[0382] This invention can significantly improve the efficiency of automated tasks while improving the user experience.
[0383] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0384] Step 1:
[0385] The user enters a request in natural language
[0386] Input: A user uses the terminal interface to type "Start the line belt every morning at 8am and collect temperature sensor data."
[0387] Data processing: The server receives the input text data and recognizes it as text data.
[0388] Output: The request text sent to the server.
[0389] Step 2:
[0390] Performing natural language analysis
[0391] Input: The request text the server received from the user.
[0392] Data processing: The server calls an external NLP engine (natural language processing engine) to analyze the request text. The prompt sent is "Text entered by the user: "Start the line belt every morning at 8am and collect temperature sensor data."
[0393] Output: Structured data as the parsed result (e.g. "Line belt started", "Temperature sensor data collected").
[0394] Step 3:
[0395] Conducting sentiment analysis
[0396] Input: The request text the server received from the user.
[0397] Data processing: The server uses an emotion analysis engine to analyze the user's emotions.
[0398] Output: Parsed emotion data (e.g. "neutral").
[0399] Step 4:
[0400] Generate RPA scripts
[0401] Input: Natural language analysis results and sentiment analysis results.
[0402] Data processing: The server generates an RPA script based on the generated structured data and emotion data. Here, the script is optimized based on the emotion data to reduce user stress.
[0403] Output: The generated RPA script.
[0404] Step 5:
[0405] Visual Indication
[0406] Input: The generated RPA script.
[0407] Data processing: Convert the script into a flowchart format and display it on the terminal.
[0408] Output: The RPA script displayed in a flowchart format.
[0409] Step 6:
[0410] User edits
[0411] Input: RPA script displayed in flowchart format.
[0412] Data processing: The user checks the flowchart on the device and edits it as necessary. The edited data is then sent back to the server.
[0413] Output: The modified RPA script.
[0414] Step 7:
[0415] Final script generation and delivery
[0416] Input: The modified RPA script.
[0417] Data processing: The server generates the final script that reflects the modifications.
[0418] Output: The final RPA script is delivered to the factory robots.
[0419] Step 8:
[0420] Robotic work execution
[0421] Input: The delivered RPA script.
[0422] Data processing: The robot performs tasks according to a script, for example, activating the line belt and collecting data from a temperature sensor.
[0423] Output: The specified result (e.g., a temperature sensor data file) is generated.
[0424] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0425] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0426] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0427] [Second embodiment]
[0428] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0429] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0430] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0431] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0432] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0433] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0434] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0435] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0436] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0437] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0438] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0439] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0440] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes an embodiment of the present invention.
[0441] The present invention relates to a system that provides an RPA tool that can be intuitively operated by a user, including an interface for users to input requests in natural language, a means for automatically generating code based on the requests, and a means for visually displaying the generated code and making it editable in a flowchart format.
[0442] Interface
[0443] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of RPA tasks in everyday languages such as English or Japanese.
[0444] Code generation
[0445] The terminal takes user input and sends a code generation request to the server, which receives the request and uses a natural language processing engine to generate Python code based on the user's requirements. The generated code is used to automatically execute RPA tasks.
[0446] The server uses an external artificial intelligence service as its natural language processing engine. This AI service analyzes natural language requests entered by users and converts them into appropriate commands and operational procedures in code. The generated code can then perform a variety of RPA tasks, such as accessing and logging into websites, retrieving data, and saving screenshots.
[0447] Code visualization and editing
[0448] The terminal visually displays the code received from the server, specifically showing the generated Python code to the user in the form of a flowchart, which illustrates how each step is processed and allows the user to intuitively understand it.
[0449] Users can check the code based on the flowchart and make any necessary modifications by clicking on a specific step in the flowchart and entering the details of the modification. For example, a user can make a modification such as "add a step to click on a specific menu after logging in."
[0450] The device then sends the user's modifications back to the server. The server then regenerates the code based on the new modifications and returns the new modified code to the device. The flowchart is updated again on the device, allowing the user to visually confirm the changes.
[0451] Specific examples
[0452] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following Python code based on the user's request (the actual content of the code is not shown here):
[0453] The generated code contains instructions to open a web browser, access the specified URL, enter a username and password to log in, take a screenshot of the dashboard page, and save it to a local file. This code is then sent back to the device and displayed as a flowchart.
[0454] The user can review the flowchart and request modifications, such as adding a step to click a specific menu after logging in. The terminal sends the modified request to the server, which then generates new code and returns it to the terminal. As a result, the RPA task required by the user can be fully executed.
[0455] The above is one embodiment of the present invention, which enables even non-programmers to intuitively create, edit, and automate RPA tasks.
[0456] The processing flow will be explained below.
[0457] Step 1:
[0458] The user accesses the system interface. The user logs in to the system using a web browser or a dedicated application.
[0459] Step 2:
[0460] The user enters a request for an RPA task into the interface in natural language, such as "Open a specific website every morning at 9am, log in, and take a screenshot of the dashboard."
[0461] Step 3:
[0462] The terminal receives the user's request and sends a code generation request to the server, which includes the user's input.
[0463] Step 4:
[0464] The server receives the request and uses a natural language processing engine to generate Python code based on the user's requirements. The server uses an external artificial intelligence service to analyze the natural language and generate the appropriate code.
[0465] Step 5:
[0466] The server sends the generated Python code in JSON format back to the device, which receives the response and parses the code.
[0467] Step 6:
[0468] It visually displays the code received by the terminal, specifically showing the generated Python code to the user in the form of a flowchart, illustrating how each step is processed.
[0469] Step 7:
[0470] The user reviews the flowchart and modifies or adds code as needed, clicking on a specific step and entering the modifications into the interface.
[0471] Step 8:
[0472] The device receives the user's correction request and sends it back to the server, containing the current code and the corrections.
[0473] Step 9:
[0474] The server receives the code generation request again and uses a natural language processing engine to generate new code based on the modifications.
[0475] Step 10:
[0476] The server generates a new code and sends it back to the device, which receives the new code and displays it again as a flowchart.
[0477] Step 11:
[0478] The user checks the updated flowchart to ensure that the changes are reflected correctly, and makes further changes if necessary.
[0479] Step 12:
[0480] Finally, once the RPA task is constructed to the user's satisfaction, the terminal executes the generated Python code, allowing the user to review the automated task and evaluate the results.
[0481] These are the program processing steps, which allows even non-programmers to intuitively create, edit, and automate RPA tasks.
[0482] Example 1
[0483] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0484] Conventional RPA tools are rarely intuitive for non-programmers to use, and require specialized knowledge, especially for generating and modifying code. Therefore, there is a need for a system that allows for input of instructions in natural language, visual code display, and editing in flowchart format. There is also a need for a system that allows for intuitive and easy modification of generated code.
[0485] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0486] In this invention, the server includes means for providing an interface for a user to input requests in natural language, means for generating code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, means for a terminal to resend user modifications to the server, the server regenerating code based on the modifications, and means for resending the resulting generated code to the terminal and updating the flowchart. This enables even non-programmers to intuitively create, edit, and automate RPA tasks.
[0487] An "interface" is a means through which a user can input requests to a system in natural language.
[0488] A "natural language request" refers to a user entering instructions or requests in everyday language.
[0489] The "code generation means" is a means having the function of automatically generating executable code based on a natural language request input by a user.
[0490] A "natural language processing engine" is an algorithm or model that analyzes natural language text, understands its content, and performs appropriate processing.
[0491] "Artificial intelligence services" refers to a general range of services that perform advanced natural language processing and machine learning using external cloud services or specific artificial intelligence models.
[0492] A "flowchart" is a diagram that visually shows each step and logic of code, and is presented in a format that is easy for users to understand intuitively.
[0493] "Terminal" refers to a device that a user uses to access and operate the system, including PCs, smartphones, tablets, etc.
[0494] A "server" is a central computer system that receives requests from users and performs code generation and other processing.
[0495] "Modifications" refers to any changes or additional instructions you make to the generated code.
[0496] A "code generation request" is a communication request to ask the server to generate code based on a user's natural language request.
[0497] This invention relates to a system that generates, modifies, and visualizes RPA task code based on user input in natural language. Specific hardware and software usage methods are described below.
[0498] Interface
[0499] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. For example, this could be a web browser such as Google® Chrome or Mozilla Firefox, or a specific application. Users can enter details of RPA tasks in everyday language such as English or Japanese.
[0500] Code generation
[0501] The terminal receives user input and sends a code generation request to the server. For example, a web application written in Python is executed through the terminal. The server receives this request and uses a natural language processing engine to generate code based on the user's requirements. The natural language processing engine includes an external generative AI model (e.g., OpenAI's GPT-3). The generated code is used to automatically perform RPA tasks, such as accessing and logging in to a website, retrieving data, and saving screenshots.
[0502] Code visualization and editing
[0503] The terminal visually displays the code received from the server. Specifically, it shows the generated Python code to the user in the form of a flowchart. The flowchart illustrates how each step is processed, allowing the user to intuitively understand it. The user can review the code based on this flowchart and make corrections as necessary. For example, they can click on a specific step in the flowchart and enter corrections. Corrections could include adding a step to click a specific menu after logging in.
[0504] Resend your fixes and update your code
[0505] The device then sends the user's modifications back to the server. The server then generates the code again based on the new modifications and returns the new modified code to the device. The flowchart is updated again on the device, allowing the user to visually confirm the changes.
[0506] Specific examples
[0507] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following Python code based on the user's request (the actual content of the code is omitted here, but the generated content is sent from the server to the terminal as is). The generated code includes steps to open a web browser, access the specified URL, enter the username and password to log in, and take a screenshot of the dashboard page and save it to a local file.
[0508] Prompt Sentence Examples
[0509] Here are some examples of prompts to input to a generative AI model:
[0510] "Every morning at 9am, open a specific website, enter your username and password, log in, take a screenshot of the dashboard, and save it to a local file."
[0511] This system allows even non-programmers to intuitively create, edit, and automate RPA tasks.
[0512] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0513] Step 1:
[0514] Users access the system's interface using a web browser or a dedicated application.
[0515] Input: User launches a web browser or dedicated application
[0516] Output: The system interface is displayed.
[0517] Specific operation: The user launches a browser on their PC or mobile device and accesses the system's URL. In the case of a dedicated application, the user connects to the interface by launching the app.
[0518] Step 2:
[0519] The user enters the details of the RPA task in natural language into input fields displayed in the interface.
[0520] Input: User inputs request in natural language
[0521] Output: A natural language request is sent to the terminal
[0522] Specific actions: The user enters information such as "Open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard," and clicks the submit button.
[0523] Step 3:
[0524] The terminal obtains the user's input and sends a code generation request to the server.
[0525] Input: A natural language request entered by the user.
[0526] Output: Code generation request to the server
[0527] Specific operation: The terminal converts the information entered by the user into a data format such as JSON, and sends the converted data to the server as an HTTP request.
[0528] Step 4:
[0529] The server performs natural language processing based on the received request and generates Python code.
[0530] Input: Code generation request sent from the device
[0531] Output: Generated Python code
[0532] How it works: The server sends a request to an external generative AI model (e.g., GPT-3), which analyzes the user's request and generates a code. The generated code is sent back to the server, which then organizes it and sends it to the device.
[0533] Step 5:
[0534] The terminal visually displays the Python code received from the server in a flowchart format.
[0535] Input: Python code sent from the server
[0536] Output: Visual representation in the form of a flowchart
[0537] Specific operation: The terminal parses the received Python code and converts each step into a flowchart node. The generated flowchart is displayed to the user.
[0538] Step 6:
[0539] The user checks the flowchart and makes corrections as necessary.
[0540] Input: Visually displayed flowchart
[0541] Output: Corrected steps or instructions
[0542] Specific operation: The user clicks on a specific step on the flowchart and enters the corrections. For example, the user can make corrections such as "add a step to click on a specific menu after logging in."
[0543] Step 7:
[0544] The terminal retransmits the user's modifications to the server, and the server generates the code again based on the modifications.
[0545] Input: The corrections entered by the user
[0546] Output: Updated Python code
[0547] Specific operation: The device converts the corrections entered by the user back into JSON format or similar and sends the converted data as an HTTP request to the server. The server then sends a new request to the AI model to generate new code. The new code is then sent back to the device, which then updates the flowchart.
[0548] In this way, even non-programmers can intuitively create, edit, and automate RPA tasks.
[0549] (Application example 1)
[0550] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0551] Inventory management at distribution centers requires a high degree of efficiency and accuracy, but often relies on manual processes and legacy systems. This can lead to problems such as delays in inventory checks and replenishment instructions, inaccurate data processing, and a lack of programming skills for automation. This results in issues such as product supply delays and excess inventory, which can reduce operational efficiency. Therefore, there is a need for automation tools that can be operated intuitively even by non-programmers.
[0552] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0553] In this invention, the server includes means for providing an interface for a user to input a request in natural language, means for generating a code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, and means for automating inventory management tasks at a logistics center based on the generated code. This makes it possible for even non-programmers to easily automate inventory management tasks and manage inventory efficiently and accurately.
[0554] "User" refers to a person who utilizes the system and enters requests in natural language.
[0555] "Natural language request" refers to the specific instructions or requests entered by the user in a natural language such as English or Japanese.
[0556] "Interface" refers to the mechanisms, such as screens and input fields, provided for users to input requests in natural language.
[0557] "Code generation means" refers to a function that automatically generates programming code based on natural language requests entered by a user.
[0558] "Flowchart format" refers to a format that visually shows the generated programming code and illustrates the flow of each step.
[0559] A "logistics center" refers to a facility that stores goods, manages inventory, processes orders, etc.
[0560] "Inventory management tasks" refer to specific inventory-related tasks such as checking inventory, issuing replenishment orders, and data management.
[0561] "Automation" refers to the automatic execution of tasks that were previously performed manually using machines or software.
[0562] A "natural language processing engine" refers to technology that analyzes natural language input by a user and generates appropriate programming code based on that.
[0563] "Artificial intelligence service" refers to a system that uses AI technology provided by external cloud services, APIs, etc. to perform specific tasks.
[0564] The present invention relates to a system for providing an inventory management automation tool for a distribution center that can be intuitively operated by a user. The system includes an interface for a user to input a request in natural language, a means for automatically generating a code based on the request, and a means for visually displaying the generated code and making it editable in a flowchart format.
[0565] Interface
[0566] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of inventory management tasks at the distribution center in everyday languages such as English and Japanese.
[0567] Code generation
[0568] The server takes user input and processes code generation requests. The server receives this request and uses a natural language processing engine to generate Python code based on the user's requirements. The generated code is used to automatically perform inventory management tasks at a distribution center. For example, it can automate tasks such as checking inventory every morning at 9:00, listing missing items, and issuing replenishment instructions.
[0569] The server uses an external artificial intelligence service as its natural language processing engine. This AI service analyzes natural language requests entered by users and converts them into appropriate operational commands and code. The generated code can then perform a variety of inventory management tasks, such as logging in, retrieving data, checking inventory, and issuing replenishment instructions.
[0570] Code visualization and editing
[0571] The terminal visually displays the code received from the server, specifically showing the generated Python code to the user in the form of a flowchart, which illustrates how each step is processed and allows the user to intuitively understand it.
[0572] The user checks the code based on the flowchart and makes any necessary corrections. Corrections are made by clicking on a specific step on the flowchart and entering the details of the correction. For example, a user could make a correction such as "adding a step that automatically issues a replenishment instruction when stock runs out." The terminal then sends the corrections back to the server. The server then generates a new code based on the new corrections and returns the new, corrected code to the terminal. The flowchart is updated again on the terminal, allowing the user to visually confirm the changes.
[0573] Specific examples
[0574] As a concrete example, consider the case where a user inputs a request such as "Check inventory every morning at 9:00, list the missing items, and issue a replenishment order." The server generates the following Python code based on the user's request (the actual content of the code is not shown here):
[0575] The generated code accesses the distribution center's database to check inventory status, lists any missing items, and includes instructions for automatically issuing replenishment instructions for those items. This code is returned to the terminal and displayed as a flowchart. The user can review the flowchart and request modifications, such as adding a step to automatically place an order with a specific supplier when inventory is low. The terminal then sends the modified request to the server, which then generates a new code and returns it to the terminal. As a result, the inventory management tasks required by the user are fully automated.
[0576] Prompt Sentence Examples
[0577] "Check inventory every morning at 9am, list any items that are in short supply, and issue replenishment instructions."
[0578] The above is an embodiment of the present invention, which enables even non-programmers to intuitively create, edit, and automate inventory management tasks.
[0579] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0580] Step 1:
[0581] The user inputs requests for inventory management tasks at the distribution center into the interface in natural language.
[0582] Input: Natural Language Request "Check inventory every morning at 9:00, list any missing items, and issue replenishment instructions."
[0583] Output: Request data to the server
[0584] Specific actions: The user uses a web browser or dedicated application to enter a request into a chat box or input field in the interface.
[0585] Step 2:
[0586] The server receives a request entered in natural language and analyzes the request using a natural language processing engine.
[0587] Input: User's natural language request
[0588] Output: Code generation request based on your requirements
[0589] Specific operation: The server makes an API call to a natural language processing engine (e.g., an external artificial intelligence service) to analyze the user's request.
[0590] Step 3:
[0591] The server generates Python code based on the analysis results and creates instructions to automate the steps of inventory management tasks.
[0592] Input: Analysis results from the natural language processing engine
[0593] Output: Python code for automated inventory management tasks
[0594] Specific operation: Based on the analysis results of the natural language processing engine, the server generates Python code to automatically perform specific inventory management tasks.
[0595] Step 4:
[0596] The server sends the generated Python code to the device.
[0597] Input: Generated Python code
[0598] Output: Data sent to the terminal
[0599] What it does: The server sends the generated Python code to the terminal and sends data for visualization in the form of a flowchart.
[0600] Step 5:
[0601] The terminal visually displays the received Python code and presents it to the user in the form of a flowchart.
[0602] Input: Python code sent from the server
[0603] Output: Visual representation in the form of a flowchart
[0604] What it does: The terminal uses a flowchart library (e.g., Graphviz) to parse the generated Python code and display each step in a flowchart format.
[0605] Step 6:
[0606] The user checks the flowchart and makes corrections as necessary.
[0607] Input: Code displayed in flowchart format
[0608] Output: User modifications
[0609] Specific operation: The user clicks on a specific step in the flowchart and enters the details of the modification, such as "add a step to automatically order from a specific supplier when inventory is low."
[0610] Step 7:
[0611] The device sends the modifications back to the server, which generates a new code.
[0612] Input: User modifications
[0613] Output: The newly generated Python code
[0614] What it does: The device sends the user's modifications to the server, and the server regenerates the Python code based on the new modifications.
[0615] Step 8:
[0616] The server sends the final modified code to the terminal, and the terminal redisplays the updated flowchart.
[0617] Input: Newly generated Python code
[0618] Output: Updated visual representation in the form of a flowchart
[0619] Specific operation: The server sends the final modified code to the terminal, and the terminal displays the updated code again as a flowchart.
[0620] Step 9:
[0621] The server runs the final generated code and automates inventory management tasks.
[0622] Input: Final revised Python code
[0623] Output: The results of the automated inventory management task.
[0624] What happens: The server executes the final generated code and automates inventory management tasks (stock checks, replenishment orders, etc.).
[0625] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0626] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes an embodiment of the present invention.
[0627] This invention relates to a system that allows users to input requests in natural language and automatically generates and edits RPA tasks. Furthermore, it aims to improve the user experience by combining it with an emotion engine that recognizes the user's emotions.
[0628] Interface
[0629] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of RPA tasks using text or voice.
[0630] Code generation
[0631] The device takes the user's input and sends a code generation request to the server. The server receives this request and uses a natural language processing engine to generate Python code based on the user's requirements. The server then uses an external artificial intelligence service to analyze the natural language and generate the appropriate code.
[0632] The server uses an external artificial intelligence service as its natural language processing engine. This AI service analyzes natural language requests entered by users and converts them into appropriate commands and operational procedures in code. The generated code can then perform a variety of RPA tasks, such as accessing and logging into websites, retrieving data, and saving screenshots.
[0633] Emotion Engine
[0634] This invention also incorporates an emotion engine, which analyzes emotions from the user's voice and text input and provides the analysis results to the natural language processing engine. This enables code generation and feedback according to the user's emotional state.
[0635] For example, if a user types "This task is tiring," the emotion engine will recognize the emotion "tiring" and provide it to the natural language processing engine, which can then use this information to more efficiently generate the appropriate code to improve the user's experience.
[0636] Code visualization and editing
[0637] The terminal visually displays the code received from the server, specifically showing the generated Python code to the user in the form of a flowchart, which illustrates how each step is processed and allows the user to intuitively understand it.
[0638] Users can review the code based on the flowchart and make corrections or additions as needed. They can click on specific steps and enter their corrections into the interface. The emotion engine also analyzes the user's emotions and provides accurate feedback.
[0639] Specific examples
[0640] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following Python code based on the user's request (the actual content of the code is not shown here):
[0641] The generated code contains instructions to open a web browser, access the specified URL, enter a username and password to log in, take a screenshot of the dashboard page, and save it to a local file. This code is then sent back to the device and displayed as a flowchart.
[0642] The user can review the flowchart and request modifications, such as adding a step to click a specific menu after logging in. The terminal sends the modified request to the server, which then generates new code and returns it to the terminal. As a result, the RPA task required by the user can be fully executed.
[0643] With the built-in emotion engine, even if the user inputs an emotional request such as "Is there a way to make this operation easier?", the system will make appropriate code modifications accordingly. In this way, the system can analyze the user's emotions and improve the user experience.
[0644] The above is one embodiment of the present invention, which enables even non-programmers to intuitively create and edit RPA tasks and automate them with an approach that takes user emotions into consideration.
[0645] The processing flow will be explained below.
[0646] Step 1:
[0647] The user accesses the system interface. The user logs in to the system using a web browser or a dedicated application.
[0648] Step 2:
[0649] The user enters a request for an RPA task into the interface in natural language, such as "Open a specific website every morning at 9am, log in, and take a screenshot of the dashboard."
[0650] Step 3:
[0651] The device receives the user's request and sends the data to the emotion engine, which analyzes the user's input text or voice to determine the emotion and sends the results back to the device.
[0652] Step 4:
[0653] Based on the emotion analysis results, the device sends a code generation request to the server. The request includes the user's input and the emotion analysis results.
[0654] Step 5:
[0655] The server receives the request and uses a natural language processing engine to generate Python code based on the user's request and sentiment analysis. The server uses an external artificial intelligence service to analyze the natural language and generate the appropriate code.
[0656] Step 6:
[0657] The server sends the generated Python code in JSON format back to the device, which receives the response and parses the code.
[0658] Step 7:
[0659] It visually displays the code received by the terminal, specifically showing the generated Python code to the user in the form of a flowchart, illustrating how each step is processed.
[0660] Step 8:
[0661] The user reviews the flowchart and makes any necessary code modifications or additions. They click on specific steps and enter their modifications into the interface. The emotion engine also analyzes the user's emotions and provides accurate feedback.
[0662] Step 9:
[0663] The device receives the user's correction request and sends it back to the server. The request includes the current code, the correction content, and the sentiment analysis results.
[0664] Step 10:
[0665] The server receives the code generation request again and uses a natural language processing engine to generate new code based on the modifications.
[0666] Step 11:
[0667] The server generates a new code and sends it back to the device, which receives the new code and displays it again as a flowchart.
[0668] Step 12:
[0669] The user checks the updated flowchart to ensure that the changes are reflected correctly, and makes further changes if necessary.
[0670] Step 13:
[0671] Finally, once the RPA task is constructed to the user's satisfaction, the terminal executes the generated Python code, allowing the user to review the automated task and evaluate the results.
[0672] These are the specific processing steps of a system incorporating an emotion engine. This enables even non-programmers to intuitively create and edit RPA tasks and automate them with an approach that takes user emotions into account.
[0673] Example 2
[0674] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0675] Conventional RPA (Robotic Process Automation) systems are often difficult for users without programming knowledge to operate intuitively. Furthermore, they do not provide an interface that takes user emotions into consideration, which does not lead to an improvement in the actual user experience. This often leaves users feeling dissatisfied and stressed.
[0676] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0677] In this invention, the server includes means for providing an interface for a user to input a request in natural language, means for generating code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, an emotion engine for analyzing the user's emotions, and means for adjusting the code based on the emotion analysis results. This allows even non-programmers to operate intuitively, and further provides feedback that takes the user's emotions into consideration, thereby improving the user experience.
[0678] An "interface" is the means by which a user accesses a system and inputs requests.
[0679] A "natural language request" is a request that a user enters into a system using natural language (e.g., everyday conversation or instructions).
[0680] The "code generation means" is a function for creating program code based on a natural language request input by a user.
[0681] "Visual display" means presenting the generated code in the form of diagrams or flowcharts so that the user can intuitively understand it.
[0682] "Editable in flowchart format" means that the user can visually check the generated code and make corrections or additions as needed.
[0683] The "emotion engine" is a function that analyzes emotions from user input (text or voice) and provides that emotional data to the system.
[0684] "Emotion analysis results" refers to the emotional information analyzed by the emotion engine from the user's input.
[0685] A "natural language processing engine" is a technology for analyzing a user's natural language requests and generating corresponding program code.
[0686] "External artificial intelligence services" means using AI technology or services provided by external providers rather than within the system.
[0687] The following describes an embodiment of the present invention. This invention relates to a system that allows users to input requests in natural language and automatically generates and edits RPA tasks. Furthermore, the invention aims to improve the user experience by combining it with an emotion engine that recognizes the user's emotions.
[0688] Interface
[0689] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of RPA tasks using text or voice.
[0690] Code generation
[0691] The device receives the user's input and sends a code generation request to the server. The server receives the request and uses a natural language processing engine to generate program code based on the user's request. The server then uses an external artificial intelligence service (e.g., GPT-4) to analyze the natural language and generate the appropriate code.
[0692] Emotion Engine
[0693] The present invention further incorporates an emotion engine. The emotion engine analyzes emotions from the user's voice or text input and provides the analysis results to the natural language processing engine. This enables code generation and feedback according to the user's emotional state. For example, if the user inputs "This task is tiring," the emotion engine recognizes the emotion "tiring" and provides it to the natural language processing engine. Based on this information, the natural language processing engine can more efficiently generate appropriate code to improve the user's experience.
[0694] Code visualization and editing
[0695] The device visually displays the code received from the server. Specifically, it shows the generated program code to the user in the form of a flowchart. The flowchart illustrates how each step is processed, allowing the user to intuitively understand. The user can review the code based on the flowchart and make corrections or additions as needed. When the user clicks on a specific step and enters the corrections into the interface, the emotion engine analyzes the user's emotions and provides accurate feedback.
[0696] Specific examples
[0697] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following program code based on the user's request (details of the code are omitted here).
[0698] The generated code contains instructions to open a web browser, access the specified URL, enter a username and password to log in, take a screenshot of the dashboard page, and save it to a local file. This code is then sent back to the device and displayed as a flowchart.
[0699] Users can review the flowchart and request modifications, such as adding a step to click a specific menu after logging in. The terminal sends the modified request to the server, which then generates new code and returns it to the terminal. As a result, the RPA task required by the user can be fully executed.
[0700] With the built-in emotion engine, even if the user inputs an emotional request such as "Is there a way to make this operation easier?", the system will make appropriate code modifications accordingly. In this way, the system can analyze the user's emotions and improve the user experience.
[0701] This invention enables even non-programmers to intuitively create and edit RPA tasks and automate them with an approach that takes user emotions into consideration.
[0702] The above is a specific embodiment for carrying out the present invention.
[0703] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0704] Step 1:
[0705] The user accesses the interface
[0706] A user accesses the system's interface using a web browser or a dedicated application. The input is the launch of the web browser or dedicated application. The output is the display of the interface, waiting for user input.
[0707] Step 2:
[0708] The user enters a request
[0709] The user inputs the details of the RPA task into the interface in natural language, either through text or voice. The input is the user's natural language request, for example, a prompt such as "Open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The output is the user's request data, which is sent to the interface.
[0710] Step 3:
[0711] The device gets user input
[0712] The terminal receives user input from the interface. As input, it receives the user's natural language request data. The terminal then prepares the request data for parallel processing. As output, it formats the request data in a well-formed way and sends it to the server.
[0713] Step 4:
[0714] The device sends a code generation request to the server
[0715] The terminal sends the user's request, properly formatted, to the server. As input, the formatted request data is sent. As output, a code generation request is issued to the server.
[0716] Step 5:
[0717] The server receives the request
[0718] The server receives a request sent from the terminal. The input is the code generation request data. The server then prepares to analyze the request data. The output is a state ready for analysis.
[0719] Step 6:
[0720] The server starts the natural language processing engine
[0721] The server launches an external AI service (e.g., GPT-4) to analyze the user's request. The input is the user's natural language request data. The output is intermediate data of the analysis results.
[0722] Step 7:
[0723] The server requests analysis from the emotion engine
[0724] The server sends the user's input data to the emotion engine and requests emotion analysis. The input is the user's input data. The output is the emotion analysis result returned from the emotion engine.
[0725] Step 8:
[0726] The emotion engine analyzes user input
[0727] The emotion engine analyzes emotions from the user's text or voice input. The input is the user's input data. Through the analysis process, the user's emotional state (e.g., "fed up") is extracted as data. The output is the emotion analysis result returned to the server.
[0728] Step 9:
[0729] The server generates code based on the analysis results
[0730] The server generates program code based on the analysis results obtained from the natural language processing engine and emotion engine. The inputs are the natural language analysis results and emotion analysis results. The output is Python code that meets the user's requirements.
[0731] Step 10:
[0732] The server sends the generated code to the device.
[0733] The server sends the generated Python code to the terminal. The input is the generated Python code. The output is the code sent to the terminal.
[0734] Step 11:
[0735] The terminal will visually display the code
[0736] The terminal visually displays the code received from the server, specifically in the form of a flowchart, to the user. The input is the Python code sent by the server. The output is a visual representation of the flowchart provided on the interface.
[0737] Step 12:
[0738] User reviews code and requests corrections
[0739] The user checks the displayed flowchart and requests corrections or additions as necessary. The inputs are the visually displayed flowchart and the user's correction requests. The output is the correction request data entered into the interface.
[0740] Step 13:
[0741] The device sends a modification request to the server
[0742] The terminal sends the user's correction request data to the server. The correction request data is the input. The output is a code generation request sent to the server again.
[0743] Step 14:
[0744] The server generates the modified code
[0745] The server generates new code based on the modification request. The input is the modification request data. The output is the updated Python code.
[0746] Step 15:
[0747] The server sends a new code to the device.
[0748] The server sends the updated Python code to the terminal. As input, it takes the updated Python code. As output, it sends the new code to the terminal.
[0749] Step 16:
[0750] The terminal will visually display the new code.
[0751] The terminal again visually displays the new code received from the server. As input, the updated Python code. As output, a visual representation of the updated flowchart is provided on the interface.
[0752] (Application example 2)
[0753] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0754] Current automation systems and robot control systems require specialized knowledge and are complex to operate, making them difficult for users to use intuitively. Furthermore, they do not take into account the user's emotions or work attitudes, resulting in problems that do not improve work efficiency or the user experience. This invention aims to provide a system that allows users to perform automated tasks by issuing instructions in natural language, and to realize a more user-friendly and efficient system by taking the user's emotions into account.
[0755] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for providing an interface for a user to input a request in natural language, means for generating code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, and emotion analysis means for analyzing the user's emotions and reflecting the analysis results in code generation. This allows the user to intuitively issue instructions in natural language, and further enables the system to generate and modify optimal code taking the user's emotional state into consideration.
[0756] An "interface" is a means by which a user can input requests in natural language.
[0757] The "code generation means" is a function that generates appropriate program code based on natural language requests entered by the user.
[0758] A "natural language processing engine" is an artificial intelligence-based system for analyzing natural language requests and generating corresponding code.
[0759] A "flowchart format" is a diagram format that visually represents each step of a program and allows users to operate it intuitively.
[0760] The "emotion analysis means" is a function that analyzes emotions from the user's voice input or text input and provides the analysis results to the system.
[0761] "External Artificial Intelligence Service" means an artificial intelligence-based service hosted on an external server that is used to analyze natural language and generate code.
[0762] A "robotic task" is a series of movements that are programmed to automate a specific task or action within a factory.
[0763] An "RPA script" is a program code written to automate software operating procedures and is used for robotic process automation.
[0764] "Emotional state" refers to the psychological state a user feels in response to a particular request or task.
[0765] The following describes the embodiments of the present invention: The system provides a set of functions for users to input requests in natural language and optimize robot operations based on sentiment analysis.
[0766] System Configuration
[0767] The system consists of a user terminal, a server, and factory robots. The user terminal provides an interface for inputting requests in natural language. The server is responsible for analyzing the natural language requests and generating the corresponding RPA script. It also analyzes the user's emotions and reflects them in the work content. Finally, the generated script is distributed to the factory robots, which then perform the specified work.
[0768] Hardware and software used
[0769] The hardware used includes users' PCs and tablets, servers (with high-performance processors and large amounts of memory), and robots that perform work in the factory.The software used includes a script generator developed in Python, an NLP (natural language processing) engine, a sentiment analysis engine, and external artificial intelligence service APIs.
[0770] Natural Language Processing and Sentiment Analysis
[0771] When a user inputs a request in natural language through their device, the request is sent to the server. The server then calls an external artificial intelligence service to analyze the input natural language. This service analyzes the user's request and generates a corresponding code. It also uses a sentiment analysis engine to analyze the user's emotions and adjust the generated code based on the results.
[0772] Visual Display and Editing
[0773] The generated code is visually displayed in a flowchart format for easy user understanding. The user can review the flowchart and edit the work steps as needed, entering modifications using the user's device interface. These modifications are then sent back to the server, where a new script is generated.
[0774] Specific examples
[0775] As a concrete example, consider the case where a user inputs "Start the line belt at 8:00 every morning and collect data from the temperature sensor." Based on this request, the system performs the following processing:
[0776] 1. A natural language processing engine analyzes the user's request and generates an RPA script to start the line belt and collect data from the temperature sensor.
[0777] 2. The sentiment analysis engine analyzes the user's emotional state and adjusts the script as needed.
[0778] 3. The generated script is displayed in flowchart format for the user to check and modify.
[0779] 4. Any modifications are processed again on the server and the final script is generated.
[0780] 5. The final script is delivered to the robot, which then performs the specified task.
[0781] This allows users to intuitively give instructions and automate tasks efficiently.
[0782] Prompt Sentence Examples
[0783] For example, if a user types "Start Linebelt every morning at 8am and collect temperature sensor data", the information passed to the system is:
[0784] User entered text: "Start Linebelt every morning at 8am and collect temperature sensor data"
[0785] This invention can significantly improve the efficiency of automated tasks while improving the user experience.
[0786] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0787] Step 1:
[0788] The user enters a request in natural language
[0789] Input: A user uses the terminal interface to type "Start the line belt every morning at 8am and collect temperature sensor data."
[0790] Data processing: The server receives the input text data and recognizes it as text data.
[0791] Output: The request text sent to the server.
[0792] Step 2:
[0793] Performing natural language analysis
[0794] Input: The request text the server received from the user.
[0795] Data processing: The server calls an external NLP engine (natural language processing engine) to analyze the request text. The prompt sent is "Text entered by the user: "Start the line belt every morning at 8am and collect temperature sensor data."
[0796] Output: Structured data as the parsed result (e.g. "Line belt started", "Temperature sensor data collected").
[0797] Step 3:
[0798] Conducting sentiment analysis
[0799] Input: The request text the server received from the user.
[0800] Data processing: The server uses an emotion analysis engine to analyze the user's emotions.
[0801] Output: Parsed emotion data (e.g. "neutral").
[0802] Step 4:
[0803] Generate RPA scripts
[0804] Input: Natural language analysis results and sentiment analysis results.
[0805] Data processing: The server generates an RPA script based on the generated structured data and emotion data. Here, the script is optimized based on the emotion data to reduce user stress.
[0806] Output: The generated RPA script.
[0807] Step 5:
[0808] Visual Indication
[0809] Input: The generated RPA script.
[0810] Data processing: Convert the script into a flowchart format and display it on the terminal.
[0811] Output: The RPA script displayed in a flowchart format.
[0812] Step 6:
[0813] User edits
[0814] Input: RPA script displayed in flowchart format.
[0815] Data processing: The user checks the flowchart on the device and edits it as necessary. The edited data is then sent back to the server.
[0816] Output: The modified RPA script.
[0817] Step 7:
[0818] Final script generation and delivery
[0819] Input: The modified RPA script.
[0820] Data processing: The server generates the final script that reflects the modifications.
[0821] Output: The final RPA script is delivered to the factory robots.
[0822] Step 8:
[0823] Robotic work execution
[0824] Input: The delivered RPA script.
[0825] Data processing: The robot performs tasks according to a script, for example, activating the line belt and collecting data from a temperature sensor.
[0826] Output: The specified result (e.g., a temperature sensor data file) is generated.
[0827] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0828] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0829] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0830] [Third embodiment]
[0831] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0832] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0833] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0834] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0835] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0836] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0837] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0838] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0839] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0840] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0841] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0842] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0843] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes an embodiment of the present invention.
[0844] The present invention relates to a system that provides an RPA tool that can be intuitively operated by a user, including an interface for users to input requests in natural language, a means for automatically generating code based on the requests, and a means for visually displaying the generated code and making it editable in a flowchart format.
[0845] Interface
[0846] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of RPA tasks in everyday languages such as English or Japanese.
[0847] Code generation
[0848] The terminal takes user input and sends a code generation request to the server, which receives the request and uses a natural language processing engine to generate Python code based on the user's requirements. The generated code is used to automatically execute RPA tasks.
[0849] The server uses an external artificial intelligence service as its natural language processing engine. This AI service analyzes natural language requests entered by users and converts them into appropriate commands and operational procedures in code. The generated code can then perform a variety of RPA tasks, such as accessing and logging into websites, retrieving data, and saving screenshots.
[0850] Code visualization and editing
[0851] The terminal visually displays the code received from the server, specifically showing the generated Python code to the user in the form of a flowchart, which illustrates how each step is processed and allows the user to intuitively understand it.
[0852] Users can check the code based on the flowchart and make any necessary modifications by clicking on a specific step in the flowchart and entering the details of the modification. For example, a user can make a modification such as "add a step to click on a specific menu after logging in."
[0853] The device then sends the user's modifications back to the server. The server then regenerates the code based on the new modifications and returns the new modified code to the device. The flowchart is updated again on the device, allowing the user to visually confirm the changes.
[0854] Specific examples
[0855] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following Python code based on the user's request (the actual content of the code is not shown here):
[0856] The generated code contains instructions to open a web browser, access the specified URL, enter a username and password to log in, take a screenshot of the dashboard page, and save it to a local file. This code is then sent back to the device and displayed as a flowchart.
[0857] The user can review the flowchart and request modifications, such as adding a step to click a specific menu after logging in. The terminal sends the modified request to the server, which then generates new code and returns it to the terminal. As a result, the RPA task required by the user can be fully executed.
[0858] The above is one embodiment of the present invention, which enables even non-programmers to intuitively create, edit, and automate RPA tasks.
[0859] The processing flow will be explained below.
[0860] Step 1:
[0861] The user accesses the system interface. The user logs in to the system using a web browser or a dedicated application.
[0862] Step 2:
[0863] The user enters a request for an RPA task into the interface in natural language, such as "Open a specific website every morning at 9am, log in, and take a screenshot of the dashboard."
[0864] Step 3:
[0865] The terminal receives the user's request and sends a code generation request to the server, which includes the user's input.
[0866] Step 4:
[0867] The server receives the request and uses a natural language processing engine to generate Python code based on the user's requirements. The server uses an external artificial intelligence service to analyze the natural language and generate the appropriate code.
[0868] Step 5:
[0869] The server sends the generated Python code in JSON format back to the device, which receives the response and parses the code.
[0870] Step 6:
[0871] It visually displays the code received by the terminal, specifically showing the generated Python code to the user in the form of a flowchart, illustrating how each step is processed.
[0872] Step 7:
[0873] The user reviews the flowchart and modifies or adds code as needed, clicking on a specific step and entering the modifications into the interface.
[0874] Step 8:
[0875] The device receives the user's correction request and sends it back to the server, containing the current code and the corrections.
[0876] Step 9:
[0877] The server receives the code generation request again and uses a natural language processing engine to generate new code based on the modifications.
[0878] Step 10:
[0879] The server generates a new code and sends it back to the device, which receives the new code and displays it again as a flowchart.
[0880] Step 11:
[0881] The user checks the updated flowchart to ensure that the changes are reflected correctly, and makes further changes if necessary.
[0882] Step 12:
[0883] Finally, once the RPA task is constructed to the user's satisfaction, the terminal executes the generated Python code, allowing the user to review the automated task and evaluate the results.
[0884] These are the program processing steps, which allows even non-programmers to intuitively create, edit, and automate RPA tasks.
[0885] Example 1
[0886] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0887] Conventional RPA tools are rarely intuitive for non-programmers to use, and require specialized knowledge, especially for generating and modifying code. Therefore, there is a need for a system that allows for input of instructions in natural language, visual code display, and editing in flowchart format. There is also a need for a system that allows for intuitive and easy modification of generated code.
[0888] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0889] In this invention, the server includes means for providing an interface for a user to input requests in natural language, means for generating code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, means for a terminal to resend user modifications to the server, the server regenerating code based on the modifications, and means for resending the resulting generated code to the terminal and updating the flowchart. This enables even non-programmers to intuitively create, edit, and automate RPA tasks.
[0890] An "interface" is a means through which a user can input requests to a system in natural language.
[0891] A "natural language request" refers to a user entering instructions or requests in everyday language.
[0892] The "code generation means" is a means having the function of automatically generating executable code based on a natural language request input by a user.
[0893] A "natural language processing engine" is an algorithm or model that analyzes natural language text, understands its content, and performs appropriate processing.
[0894] "Artificial intelligence services" refers to a general range of services that perform advanced natural language processing and machine learning using external cloud services or specific artificial intelligence models.
[0895] A "flowchart" is a diagram that visually shows each step and logic of code, and is presented in a format that is easy for users to understand intuitively.
[0896] "Terminal" refers to a device that a user uses to access and operate the system, including PCs, smartphones, tablets, etc.
[0897] A "server" is a central computer system that receives requests from users and performs code generation and other processing.
[0898] "Modifications" refers to any changes or additional instructions you make to the generated code.
[0899] A "code generation request" is a communication request to ask the server to generate code based on a user's natural language request.
[0900] This invention relates to a system that generates, modifies, and visualizes RPA task code based on user input in natural language. Specific hardware and software usage methods are described below.
[0901] Interface
[0902] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. For example, this could be a web browser such as Google Chrome or Mozilla Firefox, or a specific application. Users can enter details of RPA tasks in everyday languages such as English or Japanese.
[0903] Code generation
[0904] The terminal receives user input and sends a code generation request to the server. For example, a web application written in Python is executed through the terminal. The server receives this request and uses a natural language processing engine to generate code based on the user's requirements. The natural language processing engine includes an external generative AI model (e.g., OpenAI's GPT-3). The generated code is used to automatically perform RPA tasks, such as accessing and logging in to a website, retrieving data, and saving screenshots.
[0905] Code visualization and editing
[0906] The terminal visually displays the code received from the server. Specifically, it shows the generated Python code to the user in the form of a flowchart. The flowchart illustrates how each step is processed, allowing the user to intuitively understand it. The user can review the code based on this flowchart and make corrections as necessary. For example, they can click on a specific step in the flowchart and enter corrections. Corrections could include adding a step to click a specific menu after logging in.
[0907] Resend your fixes and update your code
[0908] The device then sends the user's modifications back to the server. The server then generates the code again based on the new modifications and returns the new modified code to the device. The flowchart is updated again on the device, allowing the user to visually confirm the changes.
[0909] Specific examples
[0910] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following Python code based on the user's request (the actual content of the code is omitted here, but the generated content is sent from the server to the terminal as is). The generated code includes steps to open a web browser, access the specified URL, enter the username and password to log in, and take a screenshot of the dashboard page and save it to a local file.
[0911] Prompt Sentence Examples
[0912] Here are some examples of prompts to input to a generative AI model:
[0913] "Every morning at 9am, open a specific website, enter your username and password, log in, take a screenshot of the dashboard, and save it to a local file."
[0914] This system allows even non-programmers to intuitively create, edit, and automate RPA tasks.
[0915] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0916] Step 1:
[0917] Users access the system's interface using a web browser or a dedicated application.
[0918] Input: User launches a web browser or dedicated application
[0919] Output: The system interface is displayed.
[0920] Specific operation: The user launches a browser on their PC or mobile device and accesses the system's URL. In the case of a dedicated application, the user connects to the interface by launching the app.
[0921] Step 2:
[0922] The user enters the details of the RPA task in natural language into input fields displayed in the interface.
[0923] Input: User inputs request in natural language
[0924] Output: A natural language request is sent to the terminal
[0925] Specific actions: The user enters information such as "Open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard," and clicks the submit button.
[0926] Step 3:
[0927] The terminal obtains the user's input and sends a code generation request to the server.
[0928] Input: A natural language request entered by the user.
[0929] Output: Code generation request to the server
[0930] Specific operation: The terminal converts the information entered by the user into a data format such as JSON, and sends the converted data to the server as an HTTP request.
[0931] Step 4:
[0932] The server performs natural language processing based on the received request and generates Python code.
[0933] Input: Code generation request sent from the device
[0934] Output: Generated Python code
[0935] How it works: The server sends a request to an external generative AI model (e.g., GPT-3), which analyzes the user's request and generates a code. The generated code is sent back to the server, which then organizes it and sends it to the device.
[0936] Step 5:
[0937] The terminal visually displays the Python code received from the server in a flowchart format.
[0938] Input: Python code sent from the server
[0939] Output: Visual representation in the form of a flowchart
[0940] Specific operation: The terminal parses the received Python code and converts each step into a flowchart node. The generated flowchart is displayed to the user.
[0941] Step 6:
[0942] The user checks the flowchart and makes corrections as necessary.
[0943] Input: Visually displayed flowchart
[0944] Output: Corrected steps or instructions
[0945] Specific operation: The user clicks on a specific step on the flowchart and enters the corrections. For example, the user can make corrections such as "add a step to click on a specific menu after logging in."
[0946] Step 7:
[0947] The terminal retransmits the user's modifications to the server, and the server generates the code again based on the modifications.
[0948] Input: The corrections entered by the user
[0949] Output: Updated Python code
[0950] Specific operation: The device converts the corrections entered by the user back into JSON format or similar and sends the converted data as an HTTP request to the server. The server then sends a new request to the AI model to generate new code. The new code is then sent back to the device, which then updates the flowchart.
[0951] In this way, even non-programmers can intuitively create, edit, and automate RPA tasks.
[0952] (Application example 1)
[0953] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0954] Inventory management at distribution centers requires a high degree of efficiency and accuracy, but often relies on manual processes and legacy systems. This can lead to problems such as delays in inventory checks and replenishment instructions, inaccurate data processing, and a lack of programming skills for automation. This results in issues such as product supply delays and excess inventory, which can reduce operational efficiency. Therefore, there is a need for automation tools that can be operated intuitively even by non-programmers.
[0955] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0956] In this invention, the server includes means for providing an interface for a user to input a request in natural language, means for generating a code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, and means for automating inventory management tasks at a logistics center based on the generated code. This makes it possible for even non-programmers to easily automate inventory management tasks and manage inventory efficiently and accurately.
[0957] "User" refers to a person who utilizes the system and enters requests in natural language.
[0958] "Natural language request" refers to the specific instructions or requests entered by the user in a natural language such as English or Japanese.
[0959] "Interface" refers to the mechanisms, such as screens and input fields, provided for users to input requests in natural language.
[0960] "Code generation means" refers to a function that automatically generates programming code based on natural language requests entered by a user.
[0961] "Flowchart format" refers to a format that visually shows the generated programming code and illustrates the flow of each step.
[0962] A "logistics center" refers to a facility that stores goods, manages inventory, processes orders, etc.
[0963] "Inventory management tasks" refer to specific inventory-related tasks such as checking inventory, issuing replenishment orders, and data management.
[0964] "Automation" refers to the automatic execution of tasks that were previously performed manually using machines or software.
[0965] A "natural language processing engine" refers to technology that analyzes natural language input by a user and generates appropriate programming code based on that.
[0966] "Artificial intelligence service" refers to a system that uses AI technology provided by external cloud services, APIs, etc. to perform specific tasks.
[0967] The present invention relates to a system for providing an inventory management automation tool for a distribution center that can be intuitively operated by a user. The system includes an interface for a user to input a request in natural language, a means for automatically generating a code based on the request, and a means for visually displaying the generated code and making it editable in a flowchart format.
[0968] Interface
[0969] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of inventory management tasks at the distribution center in everyday languages such as English and Japanese.
[0970] Code generation
[0971] The server takes user input and processes code generation requests. The server receives this request and uses a natural language processing engine to generate Python code based on the user's requirements. The generated code is used to automatically perform inventory management tasks at a distribution center. For example, it can automate tasks such as checking inventory every morning at 9:00, listing missing items, and issuing replenishment instructions.
[0972] The server uses an external artificial intelligence service as its natural language processing engine. This AI service analyzes natural language requests entered by users and converts them into appropriate operational commands and code. The generated code can then perform a variety of inventory management tasks, such as logging in, retrieving data, checking inventory, and issuing replenishment instructions.
[0973] Code visualization and editing
[0974] The terminal visually displays the code received from the server, specifically showing the generated Python code to the user in the form of a flowchart, which illustrates how each step is processed and allows the user to intuitively understand it.
[0975] The user checks the code based on the flowchart and makes any necessary corrections. Corrections are made by clicking on a specific step on the flowchart and entering the details of the correction. For example, a user could make a correction such as "adding a step that automatically issues a replenishment instruction when stock runs out." The terminal then sends the corrections back to the server. The server then generates a new code based on the new corrections and returns the new, corrected code to the terminal. The flowchart is updated again on the terminal, allowing the user to visually confirm the changes.
[0976] Specific examples
[0977] As a concrete example, consider the case where a user inputs a request such as "Check inventory every morning at 9:00, list the missing items, and issue a replenishment order." The server generates the following Python code based on the user's request (the actual content of the code is not shown here):
[0978] The generated code accesses the distribution center's database to check inventory status, lists any missing items, and includes instructions for automatically issuing replenishment instructions for those items. This code is returned to the terminal and displayed as a flowchart. The user can review the flowchart and request modifications, such as adding a step to automatically place an order with a specific supplier when inventory is low. The terminal then sends the modified request to the server, which then generates a new code and returns it to the terminal. As a result, the inventory management tasks required by the user are fully automated.
[0979] Prompt Sentence Examples
[0980] "Check inventory every morning at 9am, list any items that are in short supply, and issue replenishment instructions."
[0981] The above is an embodiment of the present invention, which enables even non-programmers to intuitively create, edit, and automate inventory management tasks.
[0982] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0983] Step 1:
[0984] The user inputs requests for inventory management tasks at the distribution center into the interface in natural language.
[0985] Input: Natural Language Request "Check inventory every morning at 9:00, list any missing items, and issue replenishment instructions."
[0986] Output: Request data to the server
[0987] Specific actions: The user uses a web browser or dedicated application to enter a request into a chat box or input field in the interface.
[0988] Step 2:
[0989] The server receives a request entered in natural language and analyzes the request using a natural language processing engine.
[0990] Input: User's natural language request
[0991] Output: Code generation request based on your requirements
[0992] Specific operation: The server makes an API call to a natural language processing engine (e.g., an external artificial intelligence service) to analyze the user's request.
[0993] Step 3:
[0994] The server generates Python code based on the analysis results and creates instructions to automate the steps of inventory management tasks.
[0995] Input: Analysis results from the natural language processing engine
[0996] Output: Python code for automated inventory management tasks
[0997] Specific operation: Based on the analysis results of the natural language processing engine, the server generates Python code to automatically perform specific inventory management tasks.
[0998] Step 4:
[0999] The server sends the generated Python code to the device.
[1000] Input: Generated Python code
[1001] Output: Data sent to the terminal
[1002] What it does: The server sends the generated Python code to the terminal and sends data for visualization in the form of a flowchart.
[1003] Step 5:
[1004] The terminal visually displays the received Python code and presents it to the user in the form of a flowchart.
[1005] Input: Python code sent from the server
[1006] Output: Visual representation in the form of a flowchart
[1007] What it does: The terminal uses a flowchart library (e.g., Graphviz) to parse the generated Python code and display each step in a flowchart format.
[1008] Step 6:
[1009] The user checks the flowchart and makes corrections as necessary.
[1010] Input: Code displayed in flowchart format
[1011] Output: User modifications
[1012] Specific operation: The user clicks on a specific step in the flowchart and enters the details of the modification, such as "add a step to automatically order from a specific supplier when inventory is low."
[1013] Step 7:
[1014] The device sends the modifications back to the server, which generates a new code.
[1015] Input: User modifications
[1016] Output: The newly generated Python code
[1017] What it does: The device sends the user's modifications to the server, and the server regenerates the Python code based on the new modifications.
[1018] Step 8:
[1019] The server sends the final modified code to the terminal, and the terminal redisplays the updated flowchart.
[1020] Input: Newly generated Python code
[1021] Output: Updated visual representation in the form of a flowchart
[1022] Specific operation: The server sends the final modified code to the terminal, and the terminal displays the updated code again as a flowchart.
[1023] Step 9:
[1024] The server runs the final generated code and automates inventory management tasks.
[1025] Input: Final revised Python code
[1026] Output: The results of the automated inventory management task.
[1027] What happens: The server executes the final generated code and automates inventory management tasks (stock checks, replenishment orders, etc.).
[1028] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1029] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes an embodiment of the present invention.
[1030] This invention relates to a system that allows users to input requests in natural language and automatically generates and edits RPA tasks. Furthermore, it aims to improve the user experience by combining it with an emotion engine that recognizes the user's emotions.
[1031] Interface
[1032] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of RPA tasks using text or voice.
[1033] Code generation
[1034] The device takes the user's input and sends a code generation request to the server. The server receives this request and uses a natural language processing engine to generate Python code based on the user's requirements. The server then uses an external artificial intelligence service to analyze the natural language and generate the appropriate code.
[1035] The server uses an external artificial intelligence service as its natural language processing engine. This AI service analyzes natural language requests entered by users and converts them into appropriate commands and operational procedures in code. The generated code can then perform a variety of RPA tasks, such as accessing and logging into websites, retrieving data, and saving screenshots.
[1036] Emotion Engine
[1037] This invention also incorporates an emotion engine, which analyzes emotions from the user's voice and text input and provides the analysis results to the natural language processing engine. This enables code generation and feedback according to the user's emotional state.
[1038] For example, if a user types "This task is tiring," the emotion engine will recognize the emotion "tiring" and provide it to the natural language processing engine, which can then use this information to more efficiently generate the appropriate code to improve the user's experience.
[1039] Code visualization and editing
[1040] The terminal visually displays the code received from the server, specifically showing the generated Python code to the user in the form of a flowchart, which illustrates how each step is processed and allows the user to intuitively understand it.
[1041] Users can review the code based on the flowchart and make corrections or additions as needed. They can click on specific steps and enter their corrections into the interface. The emotion engine also analyzes the user's emotions and provides accurate feedback.
[1042] Specific examples
[1043] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following Python code based on the user's request (the actual content of the code is not shown here):
[1044] The generated code contains instructions to open a web browser, access the specified URL, enter a username and password to log in, take a screenshot of the dashboard page, and save it to a local file. This code is then sent back to the device and displayed as a flowchart.
[1045] The user can review the flowchart and request modifications, such as adding a step to click a specific menu after logging in. The terminal sends the modified request to the server, which then generates new code and returns it to the terminal. As a result, the RPA task required by the user can be fully executed.
[1046] With the built-in emotion engine, even if the user inputs an emotional request such as "Is there a way to make this operation easier?", the system will make appropriate code modifications accordingly. In this way, the system can analyze the user's emotions and improve the user experience.
[1047] The above is one embodiment of the present invention, which enables even non-programmers to intuitively create and edit RPA tasks and automate them with an approach that takes user emotions into consideration.
[1048] The processing flow will be explained below.
[1049] Step 1:
[1050] The user accesses the system interface. The user logs in to the system using a web browser or a dedicated application.
[1051] Step 2:
[1052] The user enters a request for an RPA task into the interface in natural language, such as "Open a specific website every morning at 9am, log in, and take a screenshot of the dashboard."
[1053] Step 3:
[1054] The device receives the user's request and sends the data to the emotion engine, which analyzes the user's input text or voice to determine the emotion and sends the results back to the device.
[1055] Step 4:
[1056] Based on the emotion analysis results, the device sends a code generation request to the server. The request includes the user's input and the emotion analysis results.
[1057] Step 5:
[1058] The server receives the request and uses a natural language processing engine to generate Python code based on the user's request and sentiment analysis. The server uses an external artificial intelligence service to analyze the natural language and generate the appropriate code.
[1059] Step 6:
[1060] The server sends the generated Python code in JSON format back to the device, which receives the response and parses the code.
[1061] Step 7:
[1062] It visually displays the code received by the terminal, specifically showing the generated Python code to the user in the form of a flowchart, illustrating how each step is processed.
[1063] Step 8:
[1064] The user reviews the flowchart and makes any necessary code modifications or additions. They click on specific steps and enter their modifications into the interface. The emotion engine also analyzes the user's emotions and provides accurate feedback.
[1065] Step 9:
[1066] The device receives the user's correction request and sends it back to the server. The request includes the current code, the correction content, and the sentiment analysis results.
[1067] Step 10:
[1068] The server receives the code generation request again and uses a natural language processing engine to generate new code based on the modifications.
[1069] Step 11:
[1070] The server generates a new code and sends it back to the device, which receives the new code and displays it again as a flowchart.
[1071] Step 12:
[1072] The user checks the updated flowchart to ensure that the changes are reflected correctly, and makes further changes if necessary.
[1073] Step 13:
[1074] Finally, once the RPA task is constructed to the user's satisfaction, the terminal executes the generated Python code, allowing the user to review the automated task and evaluate the results.
[1075] These are the specific processing steps of a system incorporating an emotion engine. This enables even non-programmers to intuitively create and edit RPA tasks and automate them with an approach that takes user emotions into account.
[1076] Example 2
[1077] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1078] Conventional RPA (Robotic Process Automation) systems are often difficult for users without programming knowledge to operate intuitively. Furthermore, they do not provide an interface that takes user emotions into consideration, which does not lead to an improvement in the actual user experience. This often leaves users feeling dissatisfied and stressed.
[1079] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1080] In this invention, the server includes means for providing an interface for a user to input a request in natural language, means for generating code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, an emotion engine for analyzing the user's emotions, and means for adjusting the code based on the emotion analysis results. This allows even non-programmers to operate intuitively, and further provides feedback that takes the user's emotions into consideration, thereby improving the user experience.
[1081] An "interface" is the means by which a user accesses a system and inputs requests.
[1082] A "natural language request" is a request that a user enters into a system using natural language (e.g., everyday conversation or instructions).
[1083] The "code generation means" is a function for creating program code based on a natural language request input by a user.
[1084] "Visual display" means presenting the generated code in the form of diagrams or flowcharts so that the user can intuitively understand it.
[1085] "Editable in flowchart format" means that the user can visually check the generated code and make corrections or additions as needed.
[1086] The "emotion engine" is a function that analyzes emotions from user input (text or voice) and provides that emotional data to the system.
[1087] "Emotion analysis results" refers to the emotional information analyzed by the emotion engine from the user's input.
[1088] A "natural language processing engine" is a technology for analyzing a user's natural language requests and generating corresponding program code.
[1089] "External artificial intelligence services" means using AI technology or services provided by external providers rather than within the system.
[1090] The following describes an embodiment of the present invention. This invention relates to a system that allows users to input requests in natural language and automatically generates and edits RPA tasks. Furthermore, the invention aims to improve the user experience by combining it with an emotion engine that recognizes the user's emotions.
[1091] Interface
[1092] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of RPA tasks using text or voice.
[1093] Code generation
[1094] The device receives the user's input and sends a code generation request to the server. The server receives the request and uses a natural language processing engine to generate program code based on the user's request. The server then uses an external artificial intelligence service (e.g., GPT-4) to analyze the natural language and generate the appropriate code.
[1095] Emotion Engine
[1096] The present invention further incorporates an emotion engine. The emotion engine analyzes emotions from the user's voice or text input and provides the analysis results to the natural language processing engine. This enables code generation and feedback according to the user's emotional state. For example, if the user inputs "This task is tiring," the emotion engine recognizes the emotion "tiring" and provides it to the natural language processing engine. Based on this information, the natural language processing engine can more efficiently generate appropriate code to improve the user's experience.
[1097] Code visualization and editing
[1098] The device visually displays the code received from the server. Specifically, it shows the generated program code to the user in the form of a flowchart. The flowchart illustrates how each step is processed, allowing the user to intuitively understand. The user can review the code based on the flowchart and make corrections or additions as needed. When the user clicks on a specific step and enters the corrections into the interface, the emotion engine analyzes the user's emotions and provides accurate feedback.
[1099] Specific examples
[1100] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following program code based on the user's request (details of the code are omitted here).
[1101] The generated code contains instructions to open a web browser, access the specified URL, enter a username and password to log in, take a screenshot of the dashboard page, and save it to a local file. This code is then sent back to the device and displayed as a flowchart.
[1102] Users can review the flowchart and request modifications, such as adding a step to click a specific menu after logging in. The terminal sends the modified request to the server, which then generates new code and returns it to the terminal. As a result, the RPA task required by the user can be fully executed.
[1103] With the built-in emotion engine, even if the user inputs an emotional request such as "Is there a way to make this operation easier?", the system will make appropriate code modifications accordingly. In this way, the system can analyze the user's emotions and improve the user experience.
[1104] This invention enables even non-programmers to intuitively create and edit RPA tasks and automate them with an approach that takes user emotions into consideration.
[1105] The above is a specific embodiment for carrying out the present invention.
[1106] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1107] Step 1:
[1108] The user accesses the interface
[1109] A user accesses the system's interface using a web browser or a dedicated application. The input is the launch of the web browser or dedicated application. The output is the display of the interface, waiting for user input.
[1110] Step 2:
[1111] The user enters a request
[1112] The user inputs the details of the RPA task into the interface in natural language, either through text or voice. The input is the user's natural language request, for example, a prompt such as "Open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The output is the user's request data, which is sent to the interface.
[1113] Step 3:
[1114] The device gets user input
[1115] The terminal receives user input from the interface. As input, it receives the user's natural language request data. The terminal then prepares the request data for parallel processing. As output, it formats the request data in a well-formed way and sends it to the server.
[1116] Step 4:
[1117] The device sends a code generation request to the server
[1118] The terminal sends the user's request, properly formatted, to the server. As input, the formatted request data is sent. As output, a code generation request is issued to the server.
[1119] Step 5:
[1120] The server receives the request
[1121] The server receives a request sent from the terminal. The input is the code generation request data. The server then prepares to analyze the request data. The output is a state ready for analysis.
[1122] Step 6:
[1123] The server starts the natural language processing engine
[1124] The server launches an external AI service (e.g., GPT-4) to analyze the user's request. The input is the user's natural language request data. The output is intermediate data of the analysis results.
[1125] Step 7:
[1126] The server requests analysis from the emotion engine
[1127] The server sends the user's input data to the emotion engine and requests emotion analysis. The input is the user's input data. The output is the emotion analysis result returned from the emotion engine.
[1128] Step 8:
[1129] The emotion engine analyzes user input
[1130] The emotion engine analyzes emotions from the user's text or voice input. The input is the user's input data. Through the analysis process, the user's emotional state (e.g., "fed up") is extracted as data. The output is the emotion analysis result returned to the server.
[1131] Step 9:
[1132] The server generates code based on the analysis results
[1133] The server generates program code based on the analysis results obtained from the natural language processing engine and emotion engine. The inputs are the natural language analysis results and emotion analysis results. The output is Python code that meets the user's requirements.
[1134] Step 10:
[1135] The server sends the generated code to the device.
[1136] The server sends the generated Python code to the terminal. The input is the generated Python code. The output is the code sent to the terminal.
[1137] Step 11:
[1138] The terminal will visually display the code
[1139] The terminal visually displays the code received from the server, specifically in the form of a flowchart, to the user. The input is the Python code sent by the server. The output is a visual representation of the flowchart provided on the interface.
[1140] Step 12:
[1141] User reviews code and requests corrections
[1142] The user checks the displayed flowchart and requests corrections or additions as necessary. The inputs are the visually displayed flowchart and the user's correction requests. The output is the correction request data entered into the interface.
[1143] Step 13:
[1144] The device sends a modification request to the server
[1145] The terminal sends the user's correction request data to the server. The correction request data is the input. The output is a code generation request sent to the server again.
[1146] Step 14:
[1147] The server generates the modified code
[1148] The server generates new code based on the modification request. The input is the modification request data. The output is the updated Python code.
[1149] Step 15:
[1150] The server sends a new code to the device.
[1151] The server sends the updated Python code to the terminal. As input, it takes the updated Python code. As output, it sends the new code to the terminal.
[1152] Step 16:
[1153] The terminal will visually display the new code.
[1154] The terminal again visually displays the new code received from the server. As input, the updated Python code. As output, a visual representation of the updated flowchart is provided on the interface.
[1155] (Application example 2)
[1156] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1157] Current automation systems and robot control systems require specialized knowledge and are complex to operate, making them difficult for users to use intuitively. Furthermore, they do not take into account the user's emotions or work attitudes, resulting in problems that do not improve work efficiency or the user experience. This invention aims to provide a system that allows users to perform automated tasks by issuing instructions in natural language, and to realize a more user-friendly and efficient system by taking the user's emotions into account.
[1158] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for providing an interface for a user to input a request in natural language, means for generating code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, and emotion analysis means for analyzing the user's emotions and reflecting the analysis results in code generation. This allows the user to intuitively issue instructions in natural language, and further enables the system to generate and modify optimal code taking the user's emotional state into consideration.
[1159] An "interface" is a means by which a user can input requests in natural language.
[1160] The "code generation means" is a function that generates appropriate program code based on natural language requests entered by the user.
[1161] A "natural language processing engine" is an artificial intelligence-based system for analyzing natural language requests and generating corresponding code.
[1162] A "flowchart format" is a diagram format that visually represents each step of a program and allows users to operate it intuitively.
[1163] The "emotion analysis means" is a function that analyzes emotions from the user's voice input or text input and provides the analysis results to the system.
[1164] "External Artificial Intelligence Service" means an artificial intelligence-based service hosted on an external server that is used to analyze natural language and generate code.
[1165] A "robotic task" is a series of movements that are programmed to automate a specific task or action within a factory.
[1166] An "RPA script" is a program code written to automate software operating procedures and is used for robotic process automation.
[1167] "Emotional state" refers to the psychological state a user feels in response to a particular request or task.
[1168] The following describes the embodiments of the present invention: The system provides a set of functions for users to input requests in natural language and optimize robot operations based on sentiment analysis.
[1169] System Configuration
[1170] The system consists of a user terminal, a server, and factory robots. The user terminal provides an interface for inputting requests in natural language. The server is responsible for analyzing the natural language requests and generating the corresponding RPA script. It also analyzes the user's emotions and reflects them in the work content. Finally, the generated script is distributed to the factory robots, which then perform the specified work.
[1171] Hardware and software used
[1172] The hardware used includes users' PCs and tablets, servers (with high-performance processors and large amounts of memory), and robots that perform work in the factory.The software used includes a script generator developed in Python, an NLP (natural language processing) engine, a sentiment analysis engine, and external artificial intelligence service APIs.
[1173] Natural Language Processing and Sentiment Analysis
[1174] When a user inputs a request in natural language through their device, the request is sent to the server. The server then calls an external artificial intelligence service to analyze the input natural language. This service analyzes the user's request and generates a corresponding code. It also uses a sentiment analysis engine to analyze the user's emotions and adjust the generated code based on the results.
[1175] Visual Display and Editing
[1176] The generated code is visually displayed in a flowchart format for easy user understanding. The user can review the flowchart and edit the work steps as needed, entering modifications using the user's device interface. These modifications are then sent back to the server, where a new script is generated.
[1177] Specific examples
[1178] As a concrete example, consider the case where a user inputs "Start the line belt at 8:00 every morning and collect data from the temperature sensor." Based on this request, the system performs the following processing:
[1179] 1. A natural language processing engine analyzes the user's request and generates an RPA script to start the line belt and collect data from the temperature sensor.
[1180] 2. The sentiment analysis engine analyzes the user's emotional state and adjusts the script as needed.
[1181] 3. The generated script is displayed in flowchart format for the user to check and modify.
[1182] 4. Any modifications are processed again on the server and the final script is generated.
[1183] 5. The final script is delivered to the robot, which then performs the specified task.
[1184] This allows users to intuitively give instructions and automate tasks efficiently.
[1185] Prompt Sentence Examples
[1186] For example, if a user types "Start Linebelt every morning at 8am and collect temperature sensor data", the information passed to the system is:
[1187] User entered text: "Start Linebelt every morning at 8am and collect temperature sensor data"
[1188] This invention can significantly improve the efficiency of automated tasks while improving the user experience.
[1189] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1190] Step 1:
[1191] The user enters a request in natural language
[1192] Input: A user uses the terminal interface to type "Start the line belt every morning at 8am and collect temperature sensor data."
[1193] Data processing: The server receives the input text data and recognizes it as text data.
[1194] Output: The request text sent to the server.
[1195] Step 2:
[1196] Performing natural language analysis
[1197] Input: The request text the server received from the user.
[1198] Data processing: The server calls an external NLP engine (natural language processing engine) to analyze the request text. The prompt sent is "Text entered by the user: "Start the line belt every morning at 8am and collect temperature sensor data."
[1199] Output: Structured data as the parsed result (e.g. "Line belt started", "Temperature sensor data collected").
[1200] Step 3:
[1201] Conducting sentiment analysis
[1202] Input: The request text the server received from the user.
[1203] Data processing: The server uses an emotion analysis engine to analyze the user's emotions.
[1204] Output: Parsed emotion data (e.g. "neutral").
[1205] Step 4:
[1206] Generate RPA scripts
[1207] Input: Natural language analysis results and sentiment analysis results.
[1208] Data processing: The server generates an RPA script based on the generated structured data and emotion data. Here, the script is optimized based on the emotion data to reduce user stress.
[1209] Output: The generated RPA script.
[1210] Step 5:
[1211] Visual Indication
[1212] Input: The generated RPA script.
[1213] Data processing: Convert the script into a flowchart format and display it on the terminal.
[1214] Output: The RPA script displayed in a flowchart format.
[1215] Step 6:
[1216] User edits
[1217] Input: RPA script displayed in flowchart format.
[1218] Data processing: The user checks the flowchart on the device and edits it as necessary. The edited data is then sent back to the server.
[1219] Output: The modified RPA script.
[1220] Step 7:
[1221] Final script generation and delivery
[1222] Input: The modified RPA script.
[1223] Data processing: The server generates the final script that reflects the modifications.
[1224] Output: The final RPA script is delivered to the factory robots.
[1225] Step 8:
[1226] Robotic work execution
[1227] Input: The delivered RPA script.
[1228] Data processing: The robot performs tasks according to a script, for example, activating the line belt and collecting data from a temperature sensor.
[1229] Output: The specified result (e.g., a temperature sensor data file) is generated.
[1230] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1231] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1232] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1233] [Fourth embodiment]
[1234] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1235] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1236] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1237] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1238] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1239] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1240] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1241] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1242] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1243] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1244] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1245] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1246] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1247] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes an embodiment of the present invention.
[1248] The present invention relates to a system that provides an RPA tool that can be intuitively operated by a user, including an interface for users to input requests in natural language, a means for automatically generating code based on the requests, and a means for visually displaying the generated code and making it editable in a flowchart format.
[1249] Interface
[1250] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of RPA tasks in everyday languages such as English or Japanese.
[1251] Code generation
[1252] The terminal takes user input and sends a code generation request to the server, which receives the request and uses a natural language processing engine to generate Python code based on the user's requirements. The generated code is used to automatically execute RPA tasks.
[1253] The server uses an external artificial intelligence service as its natural language processing engine. This AI service analyzes natural language requests entered by users and converts them into appropriate commands and operational procedures in code. The generated code can then perform a variety of RPA tasks, such as accessing and logging into websites, retrieving data, and saving screenshots.
[1254] Code visualization and editing
[1255] The terminal visually displays the code received from the server, specifically showing the generated Python code to the user in the form of a flowchart, which illustrates how each step is processed and allows the user to intuitively understand it.
[1256] Users can check the code based on the flowchart and make any necessary modifications by clicking on a specific step in the flowchart and entering the details of the modification. For example, a user can make a modification such as "add a step to click on a specific menu after logging in."
[1257] The device then sends the user's modifications back to the server. The server then regenerates the code based on the new modifications and returns the new modified code to the device. The flowchart is updated again on the device, allowing the user to visually confirm the changes.
[1258] Specific examples
[1259] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following Python code based on the user's request (the actual content of the code is not shown here):
[1260] The generated code contains instructions to open a web browser, access the specified URL, enter a username and password to log in, take a screenshot of the dashboard page, and save it to a local file. This code is then sent back to the device and displayed as a flowchart.
[1261] The user can review the flowchart and request modifications, such as adding a step to click a specific menu after logging in. The terminal sends the modified request to the server, which then generates new code and returns it to the terminal. As a result, the RPA task required by the user can be fully executed.
[1262] The above is one embodiment of the present invention, which enables even non-programmers to intuitively create, edit, and automate RPA tasks.
[1263] The processing flow will be explained below.
[1264] Step 1:
[1265] The user accesses the system interface. The user logs in to the system using a web browser or a dedicated application.
[1266] Step 2:
[1267] The user enters a request for an RPA task into the interface in natural language, such as "Open a specific website every morning at 9am, log in, and take a screenshot of the dashboard."
[1268] Step 3:
[1269] The terminal receives the user's request and sends a code generation request to the server, which includes the user's input.
[1270] Step 4:
[1271] The server receives the request and uses a natural language processing engine to generate Python code based on the user's requirements. The server uses an external artificial intelligence service to analyze the natural language and generate the appropriate code.
[1272] Step 5:
[1273] The server sends the generated Python code in JSON format back to the device, which receives the response and parses the code.
[1274] Step 6:
[1275] It visually displays the code received by the terminal, specifically showing the generated Python code to the user in the form of a flowchart, illustrating how each step is processed.
[1276] Step 7:
[1277] The user reviews the flowchart and modifies or adds code as needed, clicking on a specific step and entering the modifications into the interface.
[1278] Step 8:
[1279] The device receives the user's correction request and sends it back to the server, containing the current code and the corrections.
[1280] Step 9:
[1281] The server receives the code generation request again and uses a natural language processing engine to generate new code based on the modifications.
[1282] Step 10:
[1283] The server generates a new code and sends it back to the device, which receives the new code and displays it again as a flowchart.
[1284] Step 11:
[1285] The user checks the updated flowchart to ensure that the changes are reflected correctly, and makes further changes if necessary.
[1286] Step 12:
[1287] Finally, once the RPA task is constructed to the user's satisfaction, the terminal executes the generated Python code, allowing the user to review the automated task and evaluate the results.
[1288] These are the program processing steps, which allows even non-programmers to intuitively create, edit, and automate RPA tasks.
[1289] Example 1
[1290] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1291] Conventional RPA tools are rarely intuitive for non-programmers to use, and require specialized knowledge, especially for generating and modifying code. Therefore, there is a need for a system that allows for input of instructions in natural language, visual code display, and editing in flowchart format. There is also a need for a system that allows for intuitive and easy modification of generated code.
[1292] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1293] In this invention, the server includes means for providing an interface for a user to input requests in natural language, means for generating code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, means for a terminal to resend user modifications to the server, the server regenerating code based on the modifications, and means for resending the resulting generated code to the terminal and updating the flowchart. This enables even non-programmers to intuitively create, edit, and automate RPA tasks.
[1294] An "interface" is a means through which a user can input requests to a system in natural language.
[1295] A "natural language request" refers to a user entering instructions or requests in everyday language.
[1296] The "code generation means" is a means having the function of automatically generating executable code based on a natural language request input by a user.
[1297] A "natural language processing engine" is an algorithm or model that analyzes natural language text, understands its content, and performs appropriate processing.
[1298] "Artificial intelligence services" refers to a general range of services that perform advanced natural language processing and machine learning using external cloud services or specific artificial intelligence models.
[1299] A "flowchart" is a diagram that visually shows each step and logic of code, and is presented in a format that is easy for users to understand intuitively.
[1300] "Terminal" refers to a device that a user uses to access and operate the system, including PCs, smartphones, tablets, etc.
[1301] A "server" is a central computer system that receives requests from users and performs code generation and other processing.
[1302] "Modifications" refers to any changes or additional instructions you make to the generated code.
[1303] A "code generation request" is a communication request to ask the server to generate code based on a user's natural language request.
[1304] This invention relates to a system that generates, modifies, and visualizes RPA task code based on user input in natural language. Specific hardware and software usage methods are described below.
[1305] Interface
[1306] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. For example, this could be a web browser such as Google Chrome or Mozilla Firefox, or a specific application. Users can enter details of RPA tasks in everyday languages such as English or Japanese.
[1307] Code generation
[1308] The terminal receives user input and sends a code generation request to the server. For example, a web application written in Python is executed through the terminal. The server receives this request and uses a natural language processing engine to generate code based on the user's requirements. The natural language processing engine includes an external generative AI model (e.g., OpenAI's GPT-3). The generated code is used to automatically perform RPA tasks, such as accessing and logging in to a website, retrieving data, and saving screenshots.
[1309] Code visualization and editing
[1310] The terminal visually displays the code received from the server. Specifically, it shows the generated Python code to the user in the form of a flowchart. The flowchart illustrates how each step is processed, allowing the user to intuitively understand it. The user can review the code based on this flowchart and make corrections as necessary. For example, they can click on a specific step in the flowchart and enter corrections. Corrections could include adding a step to click a specific menu after logging in.
[1311] Resend your fixes and update your code
[1312] The device then sends the user's modifications back to the server. The server then generates the code again based on the new modifications and returns the new modified code to the device. The flowchart is updated again on the device, allowing the user to visually confirm the changes.
[1313] Specific examples
[1314] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following Python code based on the user's request (the actual content of the code is omitted here, but the generated content is sent from the server to the terminal as is). The generated code includes steps to open a web browser, access the specified URL, enter the username and password to log in, and take a screenshot of the dashboard page and save it to a local file.
[1315] Prompt Sentence Examples
[1316] Here are some examples of prompts to input to a generative AI model:
[1317] "Every morning at 9am, open a specific website, enter your username and password, log in, take a screenshot of the dashboard, and save it to a local file."
[1318] This system allows even non-programmers to intuitively create, edit, and automate RPA tasks.
[1319] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1320] Step 1:
[1321] Users access the system's interface using a web browser or a dedicated application.
[1322] Input: User launches a web browser or dedicated application
[1323] Output: The system interface is displayed.
[1324] Specific operation: The user launches a browser on their PC or mobile device and accesses the system's URL. In the case of a dedicated application, the user connects to the interface by launching the app.
[1325] Step 2:
[1326] The user enters the details of the RPA task in natural language into input fields displayed in the interface.
[1327] Input: User inputs request in natural language
[1328] Output: A natural language request is sent to the terminal
[1329] Specific actions: The user enters information such as "Open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard," and clicks the submit button.
[1330] Step 3:
[1331] The terminal obtains the user's input and sends a code generation request to the server.
[1332] Input: A natural language request entered by the user.
[1333] Output: Code generation request to the server
[1334] Specific operation: The terminal converts the information entered by the user into a data format such as JSON, and sends the converted data to the server as an HTTP request.
[1335] Step 4:
[1336] The server performs natural language processing based on the received request and generates Python code.
[1337] Input: Code generation request sent from the device
[1338] Output: Generated Python code
[1339] How it works: The server sends a request to an external generative AI model (e.g., GPT-3), which analyzes the user's request and generates a code. The generated code is sent back to the server, which then organizes it and sends it to the device.
[1340] Step 5:
[1341] The terminal visually displays the Python code received from the server in a flowchart format.
[1342] Input: Python code sent from the server
[1343] Output: Visual representation in the form of a flowchart
[1344] Specific operation: The terminal parses the received Python code and converts each step into a flowchart node. The generated flowchart is displayed to the user.
[1345] Step 6:
[1346] The user checks the flowchart and makes corrections as necessary.
[1347] Input: Visually displayed flowchart
[1348] Output: Corrected steps or instructions
[1349] Specific operation: The user clicks on a specific step on the flowchart and enters the corrections. For example, the user can make corrections such as "add a step to click on a specific menu after logging in."
[1350] Step 7:
[1351] The terminal retransmits the user's modifications to the server, and the server generates the code again based on the modifications.
[1352] Input: The corrections entered by the user
[1353] Output: Updated Python code
[1354] Specific operation: The device converts the corrections entered by the user back into JSON format or similar and sends the converted data as an HTTP request to the server. The server then sends a new request to the AI model to generate new code. The new code is then sent back to the device, which then updates the flowchart.
[1355] In this way, even non-programmers can intuitively create, edit, and automate RPA tasks.
[1356] (Application example 1)
[1357] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1358] Inventory management at distribution centers requires a high degree of efficiency and accuracy, but often relies on manual processes and legacy systems. This can lead to problems such as delays in inventory checks and replenishment instructions, inaccurate data processing, and a lack of programming skills for automation. This results in issues such as product supply delays and excess inventory, which can reduce operational efficiency. Therefore, there is a need for automation tools that can be operated intuitively even by non-programmers.
[1359] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1360] In this invention, the server includes means for providing an interface for a user to input a request in natural language, means for generating a code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, and means for automating inventory management tasks at a logistics center based on the generated code. This makes it possible for even non-programmers to easily automate inventory management tasks and manage inventory efficiently and accurately.
[1361] "User" refers to a person who utilizes the system and enters requests in natural language.
[1362] "Natural language request" refers to the specific instructions or requests entered by the user in a natural language such as English or Japanese.
[1363] "Interface" refers to the mechanisms, such as screens and input fields, provided for users to input requests in natural language.
[1364] "Code generation means" refers to a function that automatically generates programming code based on natural language requests entered by a user.
[1365] "Flowchart format" refers to a format that visually shows the generated programming code and illustrates the flow of each step.
[1366] A "logistics center" refers to a facility that stores goods, manages inventory, processes orders, etc.
[1367] "Inventory management tasks" refer to specific inventory-related tasks such as checking inventory, issuing replenishment orders, and data management.
[1368] "Automation" refers to the automatic execution of tasks that were previously performed manually using machines or software.
[1369] A "natural language processing engine" refers to technology that analyzes natural language input by a user and generates appropriate programming code based on that.
[1370] "Artificial intelligence service" refers to a system that uses AI technology provided by external cloud services, APIs, etc. to perform specific tasks.
[1371] The present invention relates to a system for providing an inventory management automation tool for a distribution center that can be intuitively operated by a user. The system includes an interface for a user to input a request in natural language, a means for automatically generating a code based on the request, and a means for visually displaying the generated code and making it editable in a flowchart format.
[1372] Interface
[1373] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of inventory management tasks at the distribution center in everyday languages such as English and Japanese.
[1374] Code generation
[1375] The server takes user input and processes code generation requests. The server receives this request and uses a natural language processing engine to generate Python code based on the user's requirements. The generated code is used to automatically perform inventory management tasks at a distribution center. For example, it can automate tasks such as checking inventory every morning at 9:00, listing missing items, and issuing replenishment instructions.
[1376] The server uses an external artificial intelligence service as its natural language processing engine. This AI service analyzes natural language requests entered by users and converts them into appropriate operational commands and code. The generated code can then perform a variety of inventory management tasks, such as logging in, retrieving data, checking inventory, and issuing replenishment instructions.
[1377] Code visualization and editing
[1378] The terminal visually displays the code received from the server, specifically showing the generated Python code to the user in the form of a flowchart, which illustrates how each step is processed and allows the user to intuitively understand it.
[1379] The user checks the code based on the flowchart and makes any necessary corrections. Corrections are made by clicking on a specific step on the flowchart and entering the details of the correction. For example, a user could make a correction such as "adding a step that automatically issues a replenishment instruction when stock runs out." The terminal then sends the corrections back to the server. The server then generates a new code based on the new corrections and returns the new, corrected code to the terminal. The flowchart is updated again on the terminal, allowing the user to visually confirm the changes.
[1380] Specific examples
[1381] As a concrete example, consider the case where a user inputs a request such as "Check inventory every morning at 9:00, list the missing items, and issue a replenishment order." The server generates the following Python code based on the user's request (the actual content of the code is not shown here):
[1382] The generated code accesses the distribution center's database to check inventory status, lists any missing items, and includes instructions for automatically issuing replenishment instructions for those items. This code is returned to the terminal and displayed as a flowchart. The user can review the flowchart and request modifications, such as adding a step to automatically place an order with a specific supplier when inventory is low. The terminal then sends the modified request to the server, which then generates a new code and returns it to the terminal. As a result, the inventory management tasks required by the user are fully automated.
[1383] Prompt Sentence Examples
[1384] "Check inventory every morning at 9am, list any items that are in short supply, and issue replenishment instructions."
[1385] The above is an embodiment of the present invention, which enables even non-programmers to intuitively create, edit, and automate inventory management tasks.
[1386] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1387] Step 1:
[1388] The user inputs requests for inventory management tasks at the distribution center into the interface in natural language.
[1389] Input: Natural Language Request "Check inventory every morning at 9:00, list any missing items, and issue replenishment instructions."
[1390] Output: Request data to the server
[1391] Specific actions: The user uses a web browser or dedicated application to enter a request into a chat box or input field in the interface.
[1392] Step 2:
[1393] The server receives a request entered in natural language and analyzes the request using a natural language processing engine.
[1394] Input: User's natural language request
[1395] Output: Code generation request based on your requirements
[1396] Specific operation: The server makes an API call to a natural language processing engine (e.g., an external artificial intelligence service) to analyze the user's request.
[1397] Step 3:
[1398] The server generates Python code based on the analysis results and creates instructions to automate the steps of inventory management tasks.
[1399] Input: Analysis results from the natural language processing engine
[1400] Output: Python code for automated inventory management tasks
[1401] Specific operation: Based on the analysis results of the natural language processing engine, the server generates Python code to automatically perform specific inventory management tasks.
[1402] Step 4:
[1403] The server sends the generated Python code to the device.
[1404] Input: Generated Python code
[1405] Output: Data sent to the terminal
[1406] What it does: The server sends the generated Python code to the terminal and sends data for visualization in the form of a flowchart.
[1407] Step 5:
[1408] The terminal visually displays the received Python code and presents it to the user in the form of a flowchart.
[1409] Input: Python code sent from the server
[1410] Output: Visual representation in the form of a flowchart
[1411] What it does: The terminal uses a flowchart library (e.g., Graphviz) to parse the generated Python code and display each step in a flowchart format.
[1412] Step 6:
[1413] The user checks the flowchart and makes corrections as necessary.
[1414] Input: Code displayed in flowchart format
[1415] Output: User modifications
[1416] Specific operation: The user clicks on a specific step in the flowchart and enters the details of the modification, such as "add a step to automatically order from a specific supplier when inventory is low."
[1417] Step 7:
[1418] The device sends the modifications back to the server, which generates a new code.
[1419] Input: User modifications
[1420] Output: The newly generated Python code
[1421] What it does: The device sends the user's modifications to the server, and the server regenerates the Python code based on the new modifications.
[1422] Step 8:
[1423] The server sends the final modified code to the terminal, and the terminal redisplays the updated flowchart.
[1424] Input: Newly generated Python code
[1425] Output: Updated visual representation in the form of a flowchart
[1426] Specific operation: The server sends the final modified code to the terminal, and the terminal displays the updated code again as a flowchart.
[1427] Step 9:
[1428] The server runs the final generated code and automates inventory management tasks.
[1429] Input: Final revised Python code
[1430] Output: The results of the automated inventory management task.
[1431] What happens: The server executes the final generated code and automates inventory management tasks (stock checks, replenishment orders, etc.).
[1432] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1433] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes an embodiment of the present invention.
[1434] This invention relates to a system that allows users to input requests in natural language and automatically generates and edits RPA tasks. Furthermore, it aims to improve the user experience by combining it with an emotion engine that recognizes the user's emotions.
[1435] Interface
[1436] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of RPA tasks using text or voice.
[1437] Code generation
[1438] The device takes the user's input and sends a code generation request to the server. The server receives this request and uses a natural language processing engine to generate Python code based on the user's requirements. The server then uses an external artificial intelligence service to analyze the natural language and generate the appropriate code.
[1439] The server uses an external artificial intelligence service as its natural language processing engine. This AI service analyzes natural language requests entered by users and converts them into appropriate commands and operational procedures in code. The generated code can then perform a variety of RPA tasks, such as accessing and logging into websites, retrieving data, and saving screenshots.
[1440] Emotion Engine
[1441] This invention also incorporates an emotion engine, which analyzes emotions from the user's voice and text input and provides the analysis results to the natural language processing engine. This enables code generation and feedback according to the user's emotional state.
[1442] For example, if a user types "This task is tiring," the emotion engine will recognize the emotion "tiring" and provide it to the natural language processing engine, which can then use this information to more efficiently generate the appropriate code to improve the user's experience.
[1443] Code visualization and editing
[1444] The terminal visually displays the code received from the server, specifically showing the generated Python code to the user in the form of a flowchart, which illustrates how each step is processed and allows the user to intuitively understand it.
[1445] Users can review the code based on the flowchart and make corrections or additions as needed. They can click on specific steps and enter their corrections into the interface. The emotion engine also analyzes the user's emotions and provides accurate feedback.
[1446] Specific examples
[1447] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following Python code based on the user's request (the actual content of the code is not shown here):
[1448] The generated code contains instructions to open a web browser, access the specified URL, enter a username and password to log in, take a screenshot of the dashboard page, and save it to a local file. This code is then sent back to the device and displayed as a flowchart.
[1449] The user can review the flowchart and request modifications, such as adding a step to click a specific menu after logging in. The terminal sends the modified request to the server, which then generates new code and returns it to the terminal. As a result, the RPA task required by the user can be fully executed.
[1450] With the built-in emotion engine, even if the user inputs an emotional request such as "Is there a way to make this operation easier?", the system will make appropriate code modifications accordingly. In this way, the system can analyze the user's emotions and improve the user experience.
[1451] The above is one embodiment of the present invention, which enables even non-programmers to intuitively create and edit RPA tasks and automate them with an approach that takes user emotions into consideration.
[1452] The processing flow will be explained below.
[1453] Step 1:
[1454] The user accesses the system interface. The user logs in to the system using a web browser or a dedicated application.
[1455] Step 2:
[1456] The user enters a request for an RPA task into the interface in natural language, such as "Open a specific website every morning at 9am, log in, and take a screenshot of the dashboard."
[1457] Step 3:
[1458] The device receives the user's request and sends the data to the emotion engine, which analyzes the user's input text or voice to determine the emotion and sends the results back to the device.
[1459] Step 4:
[1460] Based on the emotion analysis results, the device sends a code generation request to the server. The request includes the user's input and the emotion analysis results.
[1461] Step 5:
[1462] The server receives the request and uses a natural language processing engine to generate Python code based on the user's request and sentiment analysis. The server uses an external artificial intelligence service to analyze the natural language and generate the appropriate code.
[1463] Step 6:
[1464] The server sends the generated Python code in JSON format back to the device, which receives the response and parses the code.
[1465] Step 7:
[1466] It visually displays the code received by the terminal, specifically showing the generated Python code to the user in the form of a flowchart, illustrating how each step is processed.
[1467] Step 8:
[1468] The user reviews the flowchart and makes any necessary code modifications or additions. They click on specific steps and enter their modifications into the interface. The emotion engine also analyzes the user's emotions and provides accurate feedback.
[1469] Step 9:
[1470] The device receives the user's correction request and sends it back to the server. The request includes the current code, the correction content, and the sentiment analysis results.
[1471] Step 10:
[1472] The server receives the code generation request again and uses a natural language processing engine to generate new code based on the modifications.
[1473] Step 11:
[1474] The server generates a new code and sends it back to the device, which receives the new code and displays it again as a flowchart.
[1475] Step 12:
[1476] The user checks the updated flowchart to ensure that the changes are reflected correctly, and makes further changes if necessary.
[1477] Step 13:
[1478] Finally, once the RPA task is constructed to the user's satisfaction, the terminal executes the generated Python code, allowing the user to review the automated task and evaluate the results.
[1479] These are the specific processing steps of a system incorporating an emotion engine. This enables even non-programmers to intuitively create and edit RPA tasks and automate them with an approach that takes user emotions into account.
[1480] Example 2
[1481] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1482] Conventional RPA (Robotic Process Automation) systems are often difficult for users without programming knowledge to operate intuitively. Furthermore, they do not provide an interface that takes user emotions into consideration, which does not lead to an improvement in the actual user experience. This often leaves users feeling dissatisfied and stressed.
[1483] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1484] In this invention, the server includes means for providing an interface for a user to input a request in natural language, means for generating code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, an emotion engine for analyzing the user's emotions, and means for adjusting the code based on the emotion analysis results. This allows even non-programmers to operate intuitively, and further provides feedback that takes the user's emotions into consideration, thereby improving the user experience.
[1485] An "interface" is the means by which a user accesses a system and inputs requests.
[1486] A "natural language request" is a request that a user enters into a system using natural language (e.g., everyday conversation or instructions).
[1487] The "code generation means" is a function for creating program code based on a natural language request input by a user.
[1488] "Visual display" means presenting the generated code in the form of diagrams or flowcharts so that the user can intuitively understand it.
[1489] "Editable in flowchart format" means that the user can visually check the generated code and make corrections or additions as needed.
[1490] The "emotion engine" is a function that analyzes emotions from user input (text or voice) and provides that emotional data to the system.
[1491] "Emotion analysis results" refers to the emotional information analyzed by the emotion engine from the user's input.
[1492] A "natural language processing engine" is a technology for analyzing a user's natural language requests and generating corresponding program code.
[1493] "External artificial intelligence services" means using AI technology or services provided by external providers rather than within the system.
[1494] The following describes an embodiment of the present invention. This invention relates to a system that allows users to input requests in natural language and automatically generates and edits RPA tasks. Furthermore, the invention aims to improve the user experience by combining it with an emotion engine that recognizes the user's emotions.
[1495] Interface
[1496] Users access the system's interface using a web browser or a dedicated application. The interface provides a chat box or input field where requests can be entered in natural language. Users can enter details of RPA tasks using text or voice.
[1497] Code generation
[1498] The device receives the user's input and sends a code generation request to the server. The server receives the request and uses a natural language processing engine to generate program code based on the user's request. The server then uses an external artificial intelligence service (e.g., GPT-4) to analyze the natural language and generate the appropriate code.
[1499] Emotion Engine
[1500] The present invention further incorporates an emotion engine. The emotion engine analyzes emotions from the user's voice or text input and provides the analysis results to the natural language processing engine. This enables code generation and feedback according to the user's emotional state. For example, if the user inputs "This task is tiring," the emotion engine recognizes the emotion "tiring" and provides it to the natural language processing engine. Based on this information, the natural language processing engine can more efficiently generate appropriate code to improve the user's experience.
[1501] Code visualization and editing
[1502] The device visually displays the code received from the server. Specifically, it shows the generated program code to the user in the form of a flowchart. The flowchart illustrates how each step is processed, allowing the user to intuitively understand. The user can review the code based on the flowchart and make corrections or additions as needed. When the user clicks on a specific step and enters the corrections into the interface, the emotion engine analyzes the user's emotions and provides accurate feedback.
[1503] Specific examples
[1504] As a concrete example, consider the case where a user inputs a request to "open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The server generates the following program code based on the user's request (details of the code are omitted here).
[1505] The generated code contains instructions to open a web browser, access the specified URL, enter a username and password to log in, take a screenshot of the dashboard page, and save it to a local file. This code is then sent back to the device and displayed as a flowchart.
[1506] Users can review the flowchart and request modifications, such as adding a step to click a specific menu after logging in. The terminal sends the modified request to the server, which then generates new code and returns it to the terminal. As a result, the RPA task required by the user can be fully executed.
[1507] With the built-in emotion engine, even if the user inputs an emotional request such as "Is there a way to make this operation easier?", the system will make appropriate code modifications accordingly. In this way, the system can analyze the user's emotions and improve the user experience.
[1508] This invention enables even non-programmers to intuitively create and edit RPA tasks and automate them with an approach that takes user emotions into consideration.
[1509] The above is a specific embodiment for carrying out the present invention.
[1510] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1511] Step 1:
[1512] The user accesses the interface
[1513] A user accesses the system's interface using a web browser or a dedicated application. The input is the launch of the web browser or dedicated application. The output is the display of the interface, waiting for user input.
[1514] Step 2:
[1515] The user enters a request
[1516] The user inputs the details of the RPA task into the interface in natural language, either through text or voice. The input is the user's natural language request, for example, a prompt such as "Open a specific website every morning at 9:00, log in, and take a screenshot of the dashboard." The output is the user's request data, which is sent to the interface.
[1517] Step 3:
[1518] The device gets user input
[1519] The terminal receives user input from the interface. As input, it receives the user's natural language request data. The terminal then prepares the request data for parallel processing. As output, it formats the request data in a well-formed way and sends it to the server.
[1520] Step 4:
[1521] The device sends a code generation request to the server
[1522] The terminal sends the user's request, properly formatted, to the server. As input, the formatted request data is sent. As output, a code generation request is issued to the server.
[1523] Step 5:
[1524] The server receives the request
[1525] The server receives a request sent from the terminal. The input is the code generation request data. The server then prepares to analyze the request data. The output is a state ready for analysis.
[1526] Step 6:
[1527] The server starts the natural language processing engine
[1528] The server launches an external AI service (e.g., GPT-4) to analyze the user's request. The input is the user's natural language request data. The output is intermediate data of the analysis results.
[1529] Step 7:
[1530] The server requests analysis from the emotion engine
[1531] The server sends the user's input data to the emotion engine and requests emotion analysis. The input is the user's input data. The output is the emotion analysis result returned from the emotion engine.
[1532] Step 8:
[1533] The emotion engine analyzes user input
[1534] The emotion engine analyzes emotions from the user's text or voice input. The input is the user's input data. Through the analysis process, the user's emotional state (e.g., "fed up") is extracted as data. The output is the emotion analysis result returned to the server.
[1535] Step 9:
[1536] The server generates code based on the analysis results
[1537] The server generates program code based on the analysis results obtained from the natural language processing engine and emotion engine. The inputs are the natural language analysis results and emotion analysis results. The output is Python code that meets the user's requirements.
[1538] Step 10:
[1539] The server sends the generated code to the device.
[1540] The server sends the generated Python code to the terminal. The input is the generated Python code. The output is the code sent to the terminal.
[1541] Step 11:
[1542] The terminal will visually display the code
[1543] The terminal visually displays the code received from the server, specifically in the form of a flowchart, to the user. The input is the Python code sent by the server. The output is a visual representation of the flowchart provided on the interface.
[1544] Step 12:
[1545] User reviews code and requests corrections
[1546] The user checks the displayed flowchart and requests corrections or additions as necessary. The inputs are the visually displayed flowchart and the user's correction requests. The output is the correction request data entered into the interface.
[1547] Step 13:
[1548] The device sends a modification request to the server
[1549] The terminal sends the user's correction request data to the server. The correction request data is the input. The output is a code generation request sent to the server again.
[1550] Step 14:
[1551] The server generates the modified code
[1552] The server generates new code based on the modification request. The input is the modification request data. The output is the updated Python code.
[1553] Step 15:
[1554] The server sends a new code to the device.
[1555] The server sends the updated Python code to the terminal. As input, it takes the updated Python code. As output, it sends the new code to the terminal.
[1556] Step 16:
[1557] The terminal will visually display the new code.
[1558] The terminal again visually displays the new code received from the server. As input, the updated Python code. As output, a visual representation of the updated flowchart is provided on the interface.
[1559] (Application example 2)
[1560] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1561] Current automation systems and robot control systems require specialized knowledge and are complex to operate, making them difficult for users to use intuitively. Furthermore, they do not take into account the user's emotions or work attitudes, resulting in problems that do not improve work efficiency or the user experience. This invention aims to provide a system that allows users to perform automated tasks by issuing instructions in natural language, and to realize a more user-friendly and efficient system by taking the user's emotions into account.
[1562] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for providing an interface for a user to input a request in natural language, means for generating code based on the natural language request input by the user, means for visually displaying the generated code and making it editable in a flowchart format, and emotion analysis means for analyzing the user's emotions and reflecting the analysis results in code generation. This allows the user to intuitively issue instructions in natural language, and further enables the system to generate and modify optimal code taking the user's emotional state into consideration.
[1563] An "interface" is a means by which a user can input requests in natural language.
[1564] The "code generation means" is a function that generates appropriate program code based on natural language requests entered by the user.
[1565] A "natural language processing engine" is an artificial intelligence-based system for analyzing natural language requests and generating corresponding code.
[1566] A "flowchart format" is a diagram format that visually represents each step of a program and allows users to operate it intuitively.
[1567] The "emotion analysis means" is a function that analyzes emotions from the user's voice input or text input and provides the analysis results to the system.
[1568] "External Artificial Intelligence Service" means an artificial intelligence-based service hosted on an external server that is used to analyze natural language and generate code.
[1569] A "robotic task" is a series of movements that are programmed to automate a specific task or action within a factory.
[1570] An "RPA script" is a program code written to automate software operating procedures and is used for robotic process automation.
[1571] "Emotional state" refers to the psychological state a user feels in response to a particular request or task.
[1572] The following describes the embodiments of the present invention: The system provides a set of functions for users to input requests in natural language and optimize robot operations based on sentiment analysis.
[1573] System Configuration
[1574] The system consists of a user terminal, a server, and factory robots. The user terminal provides an interface for inputting requests in natural language. The server is responsible for analyzing the natural language requests and generating the corresponding RPA script. It also analyzes the user's emotions and reflects them in the work content. Finally, the generated script is distributed to the factory robots, which then perform the specified work.
[1575] Hardware and software used
[1576] The hardware used includes users' PCs and tablets, servers (with high-performance processors and large amounts of memory), and robots that perform work in the factory.The software used includes a script generator developed in Python, an NLP (natural language processing) engine, a sentiment analysis engine, and external artificial intelligence service APIs.
[1577] Natural Language Processing and Sentiment Analysis
[1578] When a user inputs a request in natural language through their device, the request is sent to the server. The server then calls an external artificial intelligence service to analyze the input natural language. This service analyzes the user's request and generates a corresponding code. It also uses a sentiment analysis engine to analyze the user's emotions and adjust the generated code based on the results.
[1579] Visual Display and Editing
[1580] The generated code is visually displayed in a flowchart format for easy user understanding. The user can review the flowchart and edit the work steps as needed, entering modifications using the user's device interface. These modifications are then sent back to the server, where a new script is generated.
[1581] Specific examples
[1582] As a concrete example, consider the case where a user inputs "Start the line belt at 8:00 every morning and collect data from the temperature sensor." Based on this request, the system performs the following processing:
[1583] 1. A natural language processing engine analyzes the user's request and generates an RPA script to start the line belt and collect data from the temperature sensor.
[1584] 2. The sentiment analysis engine analyzes the user's emotional state and adjusts the script as needed.
[1585] 3. The generated script is displayed in flowchart format for the user to check and modify.
[1586] 4. Any modifications are processed again on the server and the final script is generated.
[1587] 5. The final script is delivered to the robot, which then performs the specified task.
[1588] This allows users to intuitively give instructions and automate tasks efficiently.
[1589] Prompt Sentence Examples
[1590] For example, if a user types "Start Linebelt every morning at 8am and collect temperature sensor data", the information passed to the system is:
[1591] User entered text: "Start Linebelt every morning at 8am and collect temperature sensor data"
[1592] This invention can significantly improve the efficiency of automated tasks while improving the user experience.
[1593] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1594] Step 1:
[1595] The user enters a request in natural language
[1596] Input: A user uses the terminal interface to type "Start the line belt every morning at 8am and collect temperature sensor data."
[1597] Data processing: The server receives the input text data and recognizes it as text data.
[1598] Output: The request text sent to the server.
[1599] Step 2:
[1600] Performing natural language analysis
[1601] Input: The request text the server received from the user.
[1602] Data processing: The server calls an external NLP engine (natural language processing engine) to analyze the request text. The prompt sent is "Text entered by the user: "Start the line belt every morning at 8am and collect temperature sensor data."
[1603] Output: Structured data as the parsed result (e.g. "Line belt started", "Temperature sensor data collected").
[1604] Step 3:
[1605] Conducting sentiment analysis
[1606] Input: The request text the server received from the user.
[1607] Data processing: The server uses an emotion analysis engine to analyze the user's emotions.
[1608] Output: Parsed emotion data (e.g. "neutral").
[1609] Step 4:
[1610] Generate RPA scripts
[1611] Input: Natural language analysis results and sentiment analysis results.
[1612] Data processing: The server generates an RPA script based on the generated structured data and emotion data. Here, the script is optimized based on the emotion data to reduce user stress.
[1613] Output: The generated RPA script.
[1614] Step 5:
[1615] Visual Indication
[1616] Input: The generated RPA script.
[1617] Data processing: Convert the script into a flowchart format and display it on the terminal.
[1618] Output: The RPA script displayed in a flowchart format.
[1619] Step 6:
[1620] User edits
[1621] Input: RPA script displayed in flowchart format.
[1622] Data processing: The user checks the flowchart on the device and edits it as necessary. The edited data is then sent back to the server.
[1623] Output: The modified RPA script.
[1624] Step 7:
[1625] Final script generation and delivery
[1626] Input: The modified RPA script.
[1627] Data processing: The server generates the final script that reflects the modifications.
[1628] Output: The final RPA script is delivered to the factory robots.
[1629] Step 8:
[1630] Robotic work execution
[1631] Input: The delivered RPA script.
[1632] Data processing: The robot performs tasks according to a script, for example, activating the line belt and collecting data from a temperature sensor.
[1633] Output: The specified result (e.g., a temperature sensor data file) is generated.
[1634] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1635] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1636] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1637] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1638] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1639] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1640] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1641] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1642] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1643] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1644] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1645] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1646] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1647] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1648] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1649] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1650] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1651] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1652] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1653] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1654] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1655] The following is further disclosed regarding the above embodiment.
[1656] (Claim 1)
[1657] means for providing an interface for a user to input requests in natural language;
[1658] means for generating a code based on a natural language request entered by the user;
[1659] means for visually displaying the generated code and making it editable in a flow chart format;
[1660] A system including:
[1661] (Claim 2)
[1662] 2. The system according to claim 1, wherein the code generating means automatically generates code based on a user's request using a natural language processing engine.
[1663] (Claim 3)
[1664] 10. The system of claim 1, wherein the natural language processing engine uses an external artificial intelligence service to generate the code.
[1665] "Example 1"
[1666] (Claim 1)
[1667] means for providing an interface for a user to input requests in natural language;
[1668] means for generating a code based on a natural language request entered by the user;
[1669] means for visually displaying the generated code and making it editable in a flow chart format;
[1670] a means for the terminal to retransmit the user's modifications to the server, and the server to generate a code again based on the modifications;
[1671] a means for sending the resulting code back to the terminal so that the flowchart is updated;
[1672] A system including:
[1673] (Claim 2)
[1674] 2. The system according to claim 1, wherein the code generating means automatically generates code based on a user's request using a natural language processing engine.
[1675] (Claim 3)
[1676] 10. The system of claim 1, wherein the natural language processing engine uses an external artificial intelligence service to generate the code.
[1677] "Application Example 1"
[1678] (Claim 1)
[1679] means for providing an interface for a user to input requests in natural language;
[1680] means for generating a code based on a natural language request entered by the user;
[1681] means for visually displaying the generated code and making it editable in a flow chart format;
[1682] means for automating inventory management tasks at the distribution center based on the generated codes;
[1683] A system including:
[1684] (Claim 2)
[1685] 2. The system according to claim 1, wherein the code generating means automatically generates code based on a user's request using a natural language processing engine.
[1686] (Claim 3)
[1687] 10. The system of claim 1, wherein the natural language processing engine uses an external artificial intelligence service to generate the code.
[1688] "Example 2: Combining Emotion Engines"
[1689] (Claim 1)
[1690] means for providing an interface for a user to input requests in natural language;
[1691] means for generating a code based on a natural language request entered by the user;
[1692] means for visually displaying the generated code and making it editable in a flow chart format;
[1693] an emotion engine that analyzes the emotion of the user;
[1694] means for adjusting the code based on the emotion analysis result;
[1695] A system including:
[1696] (Claim 2)
[1697] 2. The system according to claim 1, wherein the code generating means automatically generates code based on a user's request using a natural language processing engine.
[1698] (Claim 3)
[1699] 10. The system of claim 1, wherein the natural language processing engine uses an external artificial intelligence service to generate the code.
[1700] "Application example 2 when combining emotion engines"
[1701] Claims (rewritten)
[1702] (Claim 1)
[1703] means for providing an interface for a user to input requests in natural language;
[1704] means for generating a code based on a natural language request entered by the user;
[1705] means for visually displaying the generated code and making it editable in a flow chart format;
[1706] emotion analysis means for analyzing the emotion of the user and reflecting the analysis result in code generation;
[1707] A system including:
[1708] (Claim 2)
[1709] 2. The system according to claim 1, wherein the code generating means automatically generates code based on a user's request using a natural language processing engine.
[1710] (Claim 3)
[1711] 10. The system of claim 1, wherein the natural language processing engine uses an external artificial intelligence service to generate the code.
[1712] (Claim 4)
[1713] 2. The system according to claim 1, wherein said emotion analysis means changes the code generation method according to the emotional state of the user. [Explanation of symbols]
[1714] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for providing an interface for a user to input requests in natural language; means for generating a code based on a natural language request entered by the user; means for visually displaying the generated code and making it editable in a flow chart format; A system including:
2. 2. The system according to claim 1, wherein said code generating means automatically generates code based on a user's request using a natural language processing engine.
3. 10. The system of claim 1, wherein the natural language processing engine uses an external artificial intelligence service to generate the code.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A