system

A system that processes data through natural language input, queries for missing information, and combines workflows to efficiently generate data processing tasks, addressing the need for specialized tools and expertise in data analysis.

JP2026103587APending Publication Date: 2026-06-24SOFTBANK GROUP CORP

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-12
Publication Date
2026-06-24

AI Technical Summary

Technical Problem

Users often need to operate multiple specialized tools for data analysis and visualization, requiring time and effort, and lack the expertise to construct efficient workflows without specialized knowledge.

Method used

A system that accepts natural language input to understand user tasks, queries for missing information, selects and combines workflows from a template library, and provides feedback, enabling efficient data processing without specialized knowledge.

Benefits of technology

Enables users to automatically process data based on natural language input, allowing them to obtain results tailored to their needs without requiring specialized tool knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026103587000001_ABST
    Figure 2026103587000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of accepting input in natural language, A means for analyzing information from input natural language and identifying basic information related to the subject, A means of inquiring with users about missing information based on the analysis results, A means of selecting and combining appropriate work procedures from a set of templates using additional information received from the user, A means of executing the generated work procedure and providing real-time feedback of the results to the user, A means of analyzing data acquired in real time and utilizing it as optimized information for visualization in civic activities and business management, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005]

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When performing data analysis and visualization, users often need to operate multiple specialized tools, which requires time and effort to acquire and makes it difficult to efficiently construct and operate a workflow. In addition, without expertise in creating a workflow for data processing, there is a problem that the business cannot be properly carried out. The purpose of this invention is to provide an environment for efficiently performing data processing without specialized knowledge.

Means for Solving the Problems

[0005] This invention uses means to accept natural language input to understand the user's tasks and objectives, and means to analyze the input natural language to identify relevant basic information. Then, means are applied to query the user for missing information based on the analysis, and an appropriate workflow is selected and combined from a template library using the additional information provided by the user. By executing this workflow, the data processing requested by the user is realized, and by providing means to feed back the results, an efficient data processing environment that does not require specialized knowledge is provided.

[0006] "Natural language" refers to the linguistic forms that humans use on a daily basis, and is a means of expressing data and objectives in a natural way.

[0007] "Means for receiving input" refers to a function that provides an interface for users to input data processing requests in natural language.

[0008] "Means of analysis" refers to functions that interpret natural language input and extract relevant information and task details.

[0009] "Identifying basic information" means identifying the necessary information elements from the input natural language and clarifying the points for data processing.

[0010] A "means of querying missing information" refers to a function that allows users to obtain information necessary for task completion but which is not clear from the initial input.

[0011] "Additional information" refers to supplementary information that was missing from the user's instructions obtained during the analysis process.

[0012] A "template library" is a library that contains modules for efficient workflows based on past usage examples and patterns.

[0013] A "workflow" is a sequence of operations that constitutes the order in which a series of processes or tasks are executed.

[0014] A "means of providing feedback" refers to a function that communicates the results of a workflow to the user. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is a system for automatically generating data processing workflows based on natural language input from users. The system operates as follows:

[0037] First, the terminal allows users to input information in natural language through its user interface. This natural language input includes information on how the data should be processed and the required output format.

[0038] The server receives this natural language input and analyzes it using a natural language processing engine. Through this analysis, the basic information necessary for the task is extracted. For example, if a user inputs "I want to aggregate sales data by month and create a bar graph," the server identifies two main tasks: "aggregating sales data" and "creating a bar graph."

[0039] Next, the server determines whether further information is needed based on the information obtained. If necessary, the server will present specific questions to the user on the terminal. For example, "What format is the data being handled?" or "Which columns should be included in the aggregation?"

[0040] The user answers these questions and provides additional information. This process allows the server to gain a detailed understanding of the overall workflow.

[0041] The server then references a template library, selects and combines the most suitable workflow templates based on the collected information, and generates a specific workflow. This generated workflow includes steps such as data import, preprocessing, analysis, and visualization.

[0042] Finally, the server executes this workflow. The results of the workflow execution are fed back to the user in real time via the terminal. The feedback includes a success message for the execution, generated graphs, and processed data.

[0043] This system allows users to automatically process data to suit their purposes without relying on specialized tool knowledge. For example, if a marketing researcher wants to quickly grasp trends in sales data, they can simply enter their objective in natural language, and the system will automatically aggregate the data and provide a visually analyzable format.

[0044] The following describes the processing flow.

[0045] Step 1:

[0046] The device displays an interface that prompts the user for natural language input, allowing the user to input tasks and objectives related to data processing in natural language.

[0047] Step 2:

[0048] Users freely input information about their desired tasks and processes using natural language. This includes specifying the type and purpose of the data, as well as the output format.

[0049] Step 3:

[0050] The terminal receives user input and sends that data to the server.

[0051] Step 4:

[0052] The server analyzes the received natural language input using a natural language processing engine to identify basic information and key keywords necessary to accomplish the task.

[0053] Step 5:

[0054] The server identifies missing information from the analysis results and generates questions to supplement the details.

[0055] Step 6:

[0056] The server sends the generated question to the user via the terminal. The question includes data format, processing conditions, and any further details that need to be confirmed.

[0057] Step 7:

[0058] The user enters any additional information required in response to the server's questions and sends it through their terminal.

[0059] Step 8:

[0060] The server analyzes the additional information received from the user and integrates the information to ensure that all the necessary information is available to generate a complete workflow.

[0061] Step 9:

[0062] The server refers to the template library, compares it with the collected information, and selects the appropriate workflow template. Then, it combines templates as needed to build the specific workflow.

[0063] Step 10:

[0064] The server executes the generated workflow and verifies that the process runs successfully. If an error occurs during processing, it automatically identifies the problem and attempts to correct it.

[0065] Step 11:

[0066] The server sends the execution results to the terminal and provides feedback to the user. This feedback includes success messages, generated graphs, and processed data.

[0067] Step 12:

[0068] The user reviews the results sent through the device and performs further analysis and processing as needed.

[0069] This series of steps allows users to effectively and efficiently complete data-related tasks, even without specialized data processing knowledge.

[0070] (Example 1)

[0071] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0072] Data processing tasks often require specialized knowledge and skills, posing a time-consuming and cumbersome problem for non-expert users. Therefore, there is a need for systems that allow users to easily perform data processing and obtain results using natural language requests.

[0073] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0074] In this invention, the server includes a device for receiving input in natural language, a device for analyzing the input natural language and identifying information processing procedures, and a device for querying the user for additional information based on the analysis results. This allows the user to perform data processing in natural language and obtain results efficiently without requiring specialized knowledge.

[0075] A "device that accepts natural language input" is a device equipped with an interface that allows users to input data processing requests into the system using the language they use on a daily basis.

[0076] A "device that analyzes input natural language and identifies information processing procedures" is a device that uses natural language processing technology to extract the tasks and steps necessary for processing from the requests entered by the user.

[0077] A "device that prompts the user for additional information based on analysis results" is a device that interactively asks the user questions when the detailed information necessary for the task identified through analysis is insufficient.

[0078] A "device that selects and combines processing procedures to generate output based on additional information obtained from the user" is a device that utilizes the user's response to select the optimal processing procedure from a template library and combines them to construct a series of data processing tasks.

[0079] A "device that executes a generated processing procedure and provides the execution result to the user via a display device" is a device that executes an assembled processing procedure and provides the result to the user in a form that is visualized or displayed.

[0080] A "generative artificial intelligence model" is a type of artificial intelligence that has the function of analyzing natural language and generating information, and is a technology that enables advanced language understanding and response generation.

[0081] An "interactive user interface" is an interface that allows the user and the system to exchange information, ask appropriate questions and give instructions, and complete tasks sequentially.

[0082] This invention is an information processing system that enables users to perform data processing using natural language without requiring specialized knowledge. In implementing this system, analysis using natural language processing technology and an interactive interface play important roles.

[0083] User: Users access the system through a terminal and input data processing requests in natural language. For example, they can input a specific command such as, "I want to aggregate sales data by month and create a bar graph."

[0084] Terminal: The terminal receives natural language input from the user and provides a user interface that can communicate with the database. This allows the user to review their input and prepare it for transmission to the server. The interactive interface on the terminal can also flexibly accept additional user input.

[0085] Server: The server uses natural language input received from the terminal to analyze the input data using generative AI models (e.g., GPT-3®, BERT). Through this analysis, task information necessary for data processing is extracted. Furthermore, the server refers to a template library to select appropriate processing steps and automatically generates an optimal workflow based on the user's requirements.

[0086] Specific example of operation: For instance, a marketing person inputs into the system, "I want to visualize this month's sales data in a bar graph." Based on this request, the server analyzes and applies the sales data processing procedure, and as a result generates the visualized graph.

[0087] By utilizing a generative AI model and prompt messages, processing results are generated in real time and feedback is provided to the user via the terminal. This system allows users to experience the automation and efficiency improvements of data processing without a steep learning curve.

[0088] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0089] Step 1:

[0090] User: The user inputs data processing requests into the system via a terminal using natural language. This input includes specifying how the data should be aggregated and the output format. An example of input is the text, "I want to aggregate sales data by month and create a bar graph."

[0091] Step 2:

[0092] Terminal: The terminal receives natural language input from the user and sends it to the server. It provides a UI to prompt the user to confirm the input, and after the user confirms the input, it prepares to send it to the server. The input here is the natural language request from step 1, and the output is the data sent to the server.

[0093] Step 3:

[0094] Server: The server uses a natural language processing engine to analyze input data from the terminal. It utilizes a generative AI model to decompose and analyze the input text into tokens to identify data processing tasks. The input is natural language text from the terminal, and the output is an internal data structure that defines the task.

[0095] Step 4:

[0096] Server: After the task is identified, the server determines the missing information and generates additional questions as needed. These questions are designed to gather detailed information about the target data format and the data to be aggregated. The input is the analysis result, and the output is the additional question text.

[0097] Step 5:

[0098] Terminal: Displays additional questions from the server to the user and provides an interactive interface for the user to answer. Input is the questions from the server, and output is the user's answer.

[0099] Step 6:

[0100] User: The user enters specific answers to additional questions via their terminal. For example, they might answer, "I want to use data in CSV format and aggregate the year and month columns." The input is the server's question, and the output is the user's additional information.

[0101] Step 7:

[0102] Server: Receives additional user information, references a template library to select an appropriate workflow, and generates specific processing steps. Input is the user's additional information, and output is the automatically generated workflow.

[0103] Step 8:

[0104] Server: Executes the generated workflow and compiles the processing results. Specific operations include data import, preprocessing, analysis, and visualization. The input is the workflow, and the output is the processing result.

[0105] Step 9:

[0106] Terminal: Receives processing results from the server and provides feedback to the user. This includes visualized results such as bar graphs. Input is the processing result from the server, and output is the feedback displayed to the user.

[0107] This process allows users to automatically process data based on simple natural language input and obtain results tailored to their specific needs.

[0108] (Application Example 1)

[0109] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0110] Efficiently processing large amounts of data in cities and providing real-time information useful to citizens and the administration is crucial for optimizing urban management. However, conventional data processing systems require specialized knowledge and make real-time information acquisition and visualization difficult. Furthermore, advanced technical skills were necessary for users to perform data analysis based on their individual needs.

[0111] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0112] In this invention, the server includes means for receiving input in natural language, means for analyzing information from the input natural language to identify basic information, and means for analyzing data acquired in real time and utilizing it as optimized information for visualization in urban activities and business management. As a result, users can perform real-time data processing and analysis of urban data and obtain optimized information by giving instructions in natural language without requiring specialized knowledge.

[0113] "Natural language input" refers to instructions or requests that users make to a computer system using the language that humans use in everyday life.

[0114] "Analysis" is the process of breaking down input natural language data, understanding the information, and extracting the necessary data.

[0115] "Basic information" refers to the data obtained through analysis and the key data and conditions necessary to execute the instructions.

[0116] A "template collection" is a library containing templates for various processing procedures and workflows, which are selected according to specific needs.

[0117] A "work procedure" is a series of processing steps necessary to achieve a specific goal.

[0118] "Real-time feedback" refers to a state where the results of a process are immediately returned to the user, allowing for immediate confirmation and correction.

[0119] "Using it for civic activities and business management" means applying the analyzed data and information in a way that is useful for citizens' lives and administrative operations.

[0120] "Generative artificial intelligence" is a form of artificial intelligence technology that analyzes human natural language and generates flows for instructions and data processing.

[0121] An "interactive interface" refers to a screen or method that allows users to interact with a system and supplement information.

[0122] "Optimizing urban operations" refers to initiatives aimed at efficiently and effectively managing and operating resources and activities within a city.

[0123] To implement this invention, a data processing system based on natural language input is required. The main components of the system include a terminal for receiving natural language input from the user and a server that analyzes the input data and generates the optimal work procedure.

[0124] First, the terminal accepts natural language input from the user via voice or text. This process is carried out via a smartphone or smart glasses, providing an environment where the user can easily input instructions. Next, the server analyzes the input using natural language processing technologies such as Google® Cloud Natural Language API and extracts the necessary information for data processing. This analysis clarifies specific tasks, such as updating traffic data or visualizing heatmaps.

[0125] Once the analysis is complete, the server selects the optimal procedure from a set of templates and performs the data processing necessary for city operation and management in real time. The generated results are returned to the terminal in an easy-to-understand format using data visualization libraries such as D3.js and Plotly. This allows users to quickly review the information and take appropriate action.

[0126] As a concrete example, if a user prompts the system with a message such as, "Analyze the air quality data in the city and identify areas that need improvement," the system will immediately acquire the data and generate information useful for urban management. In this way, the invention contributes to optimizing urban operations.

[0127] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0128] Step 1:

[0129] Users input instructions in natural language. For example, a user might use a smartphone or smart glasses to say, "Update the traffic data and display congested areas on a heatmap." At this stage, natural language voice data or text data is obtained as input.

[0130] Step 2:

[0131] The terminal converts received natural language input into text data and sends it to the server. In the case of voice input, speech recognition technology is used to convert it into text. The output here is an instruction in a parseable text format.

[0132] Step 3:

[0133] The server uses the Google Cloud Natural Language API to analyze this text data and extract the necessary tasks and data from the instructions. This process identifies the analyzed instructions and clarifies tasks such as "update traffic data" and "display heatmaps of congested areas." This is the output from the server.

[0134] Step 4:

[0135] Based on the analysis results, the server selects the optimal work procedure from a set of templates and constructs it. At this stage, predefined processing steps within the template library are referenced to select the necessary data acquisition methods and processing steps. The output is the generated workflow.

[0136] Step 5:

[0137] The server retrieves and processes data according to the created workflow. For example, it retrieves the latest traffic data from an API as needed and analyzes it in the specified format. At this point, calculations are performed based on the input data, and data suitable for visualization is generated as output.

[0138] Step 6:

[0139] The server uses visualization libraries such as D3.js and Plotly to visualize the generated data in a user-friendly format. The created heatmaps and analysis results are then fed back to the user's device.

[0140] Step 7:

[0141] The terminal displays real-time feedback from the server to the user. Traffic data is updated as instructed by the user and visually displayed as a heatmap. This allows the user to quickly obtain the necessary information and utilize it.

[0142] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0143] This invention combines a system that automatically generates data processing workflows using natural language input from users with an emotion engine that recognizes user emotions. This system is implemented in the following way.

[0144] First, the device provides a platform for the user to input natural language information through its user interface. This input includes specific purposes and desires regarding data manipulation.

[0145] The server receives natural language input from the terminal and analyzes it using a natural language processing engine. Basic information regarding data processing is extracted here. Simultaneously, the emotion engine is activated to recognize underlying emotions from the user's input. This emotion information is used to adjust the system's response and processing methods.

[0146] Next, the server determines, based on the analysis results, whether any information is missing. If necessary, it also takes into account the emotional state and generates questions to ask the user for additional information in an appropriate tone and expression. These questions are then displayed to the user again through the terminal.

[0147] The user answers questions displayed on the device and provides any necessary additional information. The user's responses are also monitored by an emotion engine. The server integrates this information and generates an appropriate workflow by selecting or combining elements from a template library.

[0148] The generated workflow is executed by the server, and the user's emotions during the process are fed back into the system, dynamically adjusting the workflow. The final processing result is notified to the user via the terminal. The method of presenting this result is also customized according to the user's emotions.

[0149] For example, when a user analyzes sales data, if the system does not meet the user's expectations, it senses this emotion and provides more satisfying results by suggesting different data perspectives or adding process visualizations. Furthermore, when making improvement suggestions, the system takes the user's stress level into consideration and adjusts the information to be presented in a way that is easy for them to accept.

[0150] In this way, we can provide an interactive data processing environment that is adapted to the user's needs and emotions.

[0151] The following describes the processing flow.

[0152] Step 1:

[0153] The terminal provides users with an interface that accepts natural language input, allowing them to freely input their data processing objectives and preferences.

[0154] Step 2:

[0155] The user inputs and sends data processing instructions in natural language via their device, for example, "I want to create a monthly report and analyze trends."

[0156] Step 3:

[0157] The server uses a natural language processing engine to analyze the natural language input it receives and extracts basic information to construct the data processing task.

[0158] Step 4:

[0159] The server simultaneously uses an emotion engine to analyze the user's emotional state based on natural language input and recognize underlying emotions.

[0160] Step 5:

[0161] Based on the basic information and emotional state identified from the analysis results, the server identifies missing information and generates questions to query the user for that information. The tone and content of the questions are adjusted according to the user's emotional state.

[0162] Step 6:

[0163] The server sends the generated question to the terminal and requests additional information from the user.

[0164] Step 7:

[0165] The user enters any necessary additional information in response to the questions displayed on the device and sends it to the server via the device.

[0166] Step 8:

[0167] The server receives additional information provided by the user, refers to a template library to select an appropriate data processing workflow, and then combines and generates it.

[0168] Step 9:

[0169] The server executes the generated workflow while taking the user's emotional state into consideration, and monitors the process to ensure that no problems occur as it progresses.

[0170] Step 10:

[0171] The server collects the execution results, adjusts them according to the user's emotions, and sends them to the terminal for feedback.

[0172] Step 11:

[0173] The device displays feedback sent from the server to the user, presenting it in a format that makes it easy for the user to understand the results.

[0174] This series of steps, through a process that takes user emotions into account, provides more appropriate and flexible data processing results.

[0175] (Example 2)

[0176] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0177] Current data processing systems can adequately analyze users' natural language input to generate optimal data processing workflows, but they struggle to respond to changes in user emotions and incorporate that feedback. As a result, users may have a less satisfying experience, and the system fails to meet expectations for interactive data processing.

[0178] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0179] In this invention, the server includes means for analyzing natural language input, means for recognizing user emotions, means for adjusting responses based on emotions, and means for selecting and generating workflows from a template library. This enables real-time responses to both user input and emotions, and provides optimized data processing workflows.

[0180] "Means of accepting input in natural language" refers to an interface that allows users to provide instructions and information to a system using natural language.

[0181] "Means of analyzing information and identifying basic information related to the subject" refers to the process of analyzing natural language input provided by the user and extracting basic information necessary for data processing from it.

[0182] "Means of querying users for missing information" refers to a mechanism that requests the user, in the form of a question, to complete the missing information identified as a result of the analysis.

[0183] "Means of understanding user emotions" refers to technologies that recognize emotional states from users' natural language input and responses.

[0184] "Means of adjusting response methods based on emotional information" refers to a process that modifies the system's responses and presentation methods according to the recognized emotions of the user, in order to provide an appropriate experience for the user.

[0185] "A means of selecting and combining appropriate workflows from a template library" refers to a method of constructing specific processing procedures by selecting the template that best suits the user's requirements from existing templates, or by combining multiple templates.

[0186] "A means of executing the generated workflow and providing feedback to the user" refers to the process of actually running the constructed data processing workflow and notifying the user of the processing results.

[0187] The system for carrying out this invention operates on a computer network including servers, terminals, and a user interface. Some processing utilizes a cloud-based platform.

[0188] The device provides a user interface, creating a space for users to input information using natural language. This interface allows input via text boxes or voice input, and utilizes tools such as the Google Cloud Natural Language API for natural language processing.

[0189] The server receives natural language input data sent from the terminal and analyzes the information via a natural language processing engine. This engine meticulously analyzes the input sentences and accurately extracts the basic information necessary for data processing. Simultaneously, the server uses sentiment recognition technology to understand the user's emotional state in real time from their input. Sentiment analysis utilizes technologies such as Microsoft® Azure® Text Analytics.

[0190] Based on emotional information, the server generates questions in an appropriate tone and content to ask the user for missing information, and presents them to the user through the terminal. By utilizing a generative AI model to adjust the prompt text as needed, the quality of communication is improved.

[0191] The user can respond to questions presented on the device and provide any necessary additional information. The emotion recognition engine also operates during this process, providing appropriate feedback based on the user's responses.

[0192] The server integrates all information obtained from the user and selects or generates the most appropriate workflow from the template library. This workflow is dynamically adjusted to match the user's needs based on the generated AI model.

[0193] The processing results obtained from the executed workflow are customized to the user's emotions and ultimately fed back to the user via the device. The results may also be displayed as infographics, providing a visually easy-to-understand format.

[0194] For example, if a user requests, "Find outliers in recent sales data," the system will analyze the input data and ask additional questions about the required data range and calculation criteria. Based on the user's response, it will generate and execute an appropriate workflow and provide feedback on the results in an appropriate format.

[0195] A concrete example of a prompt statement is, "Please tell me the data range required for this process." This allows for the provision of information that aligns with the user's needs and emotions.

[0196] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0197] Step 1:

[0198] The device provides a user interface, allowing users to input information and instructions in natural language. Users can use either a text box or voice input. The input data is captured by the device as natural language text.

[0199] Step 2:

[0200] The server receives natural language text from the terminal as input and sends it to the natural language processing engine. The natural language processing engine analyzes the text and extracts the basic information necessary for data processing. This process includes identifying key keywords and context.

[0201] Step 3:

[0202] Based on the analyzed information, the server uses an emotion recognition engine to recognize the user's emotional state. This input includes natural language text, and the output is either an emotion score or an emotion category. This allows for an understanding of the user's emotions.

[0203] Step 4:

[0204] The server identifies the missing information and generates a question in an appropriate tone, taking sentiment scores into consideration. Using a generative AI model, it creates a prompt and presents a specific question (e.g., "Please tell me the data range required for this process"). This question is then sent to the user.

[0205] Step 5:

[0206] The user enters their answers to the presented questions into the terminal. The additional information entered is specific response data to the questions. This allows for data completion.

[0207] Step 6:

[0208] The server retrieves additional information from the user and selects the most suitable workflow from the template library. Based on the retrieved additional information, data calculations and processing are performed, and the workflow is generated.

[0209] Step 7:

[0210] The server executes the generated workflow. This process performs each step for data processing, and in the process, user sentiment data is re-evaluated. The execution results are then obtained.

[0211] Step 8:

[0212] The server sends the final processing results to the terminal and notifies the user. The results are presented in a customized format based on the user's emotions. This output may take the form of an infographic or summary.

[0213] (Application Example 2)

[0214] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0215] Conventional natural language processing systems do not consider user emotions when analyzing input, and therefore do not optimize responses based on emotions. This can lead to decreased user satisfaction. Furthermore, fixed dialogue methods lack flexibility in sequential information supplementation, making it difficult to provide dynamic information that meets user needs.

[0216] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0217] In this invention, the server includes means for receiving input in natural language, means for analyzing information from the input natural language and identifying basic information related to the subject, and means for recognizing the user's emotions and adjusting the response content according to those emotions. This enables responses that are adapted to the user's emotions, improving the quality of interaction.

[0218] A "means of accepting natural language input" refers to an interface for receiving data and instructions from users in the form of natural language.

[0219] "Means for analyzing information and identifying basic information related to a subject" refers to the process of analyzing the content of natural language received from a user and extracting key points and basic information related to a specific subject based on that content.

[0220] "Means of querying users for missing information" refers to methods of identifying shortcomings in user instructions based on analysis results and asking users about them.

[0221] "Means of recognizing user emotions and adjusting responses accordingly" refers to a mechanism that analyzes the user's emotional state and dynamically changes the content of responses and information presented based on the results.

[0222] "A means of selecting and combining appropriate work procedures from a template library" refers to a system for selecting the most suitable predefined work procedures according to the situation and combining them as needed to construct new procedures.

[0223] "A means of executing the generated work procedure and providing feedback to the user on the results" refers to a function that actually carries out the created work procedure and reports the results and progress of that work to the user.

[0224] This invention relates to a system applied to a consumer robot that provides household support by utilizing natural language input from the user. The system components include a natural language processing engine (NLP engine), an emotion recognition engine, and a task scheduler. The hardware is a home robot equipped with a microphone for receiving voice input, a speaker for outputting responses, and a display for providing visual information.

[0225] The server first receives natural language input from the user through the robot's microphone. The received audio data is analyzed using an NLP engine to understand the user's request. At the same time, the emotion recognition engine is activated to recognize the user's emotional state in real time. This information is used to adjust the response in subsequent processes.

[0226] Once the analysis is complete, the server selects the optimal household chore procedure from the template library and, if necessary, combines and adjusts the order of tasks in a way that adapts to the user's emotions. The created procedure is then executed by a robot, and the progress and results are fed back to the user. The feedback re-evaluates the user's emotions and is reported in an appropriate format that reflects those emotions.

[0227] As a concrete example, suppose a user instructs a robot to "clean the kitchen next Friday." The NLP engine understands this instruction, and the emotion recognition engine detects that the user is tired. In this case, the robot proposes a more efficient cleaning plan and schedules the task in a way that reduces the user's burden.

[0228] Examples of prompts include, "How would the robot respond if the user said, 'When I've finished cleaning the kitchen, suggest a new recipe?'" and "How would you improve the suggestion if the user seemed dissatisfied?" In this way, the system enables efficient and adaptive responses based on the user's instructions and emotions.

[0229] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0230] Step 1:

[0231] The user inputs natural language instructions to the robot via a microphone. This input is voice data, which the terminal collects. This prepares the voice data for transmission to the server.

[0232] Step 2:

[0233] The server provides the audio data received from the terminal to the natural language processing engine, where it is converted into text-based instructions. This text data includes information about specific household tasks and desired actions. This conversion enables analysis in the next step.

[0234] Step 3:

[0235] The server uses the analysis capabilities of its natural language processing engine to extract important information from the text data. It identifies keywords and tasks related to the subject, and then uses an emotion recognition engine to evaluate the user's emotional state. At this stage, input data (text) and intermediate data (emotional information) are generated.

[0236] Step 4:

[0237] The server generates questions to prompt the user for additional information if any is missing. It adjusts the questions, including the appropriate tone and phrasing, taking sentiment recognition results into account. These questions are then sent to the terminal and displayed to the user, setting the stage for obtaining the user's response in the next step.

[0238] Step 5:

[0239] The user provides additional information in response to questions displayed on the device. This response is again collected as audio data and sent to the server. The user's emotional state is also monitored simultaneously.

[0240] Step 6:

[0241] The server integrates additional information and sentiment data received from the user, selects the optimal work procedure from the template library, and combines it as needed to generate a new workflow. This generated workflow is then ready to be executed by the task scheduler.

[0242] Step 7:

[0243] A workflow built by the server is executed by a robot. During the process, the user's emotions are continuously evaluated, and the workflow is dynamically adjusted based on that feedback. Finally, the results are reported to the user via a terminal. This report is presented in a format that reflects the user's emotions.

[0244] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0245] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0246] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0247] [Second Embodiment]

[0248] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0249] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0250] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0251] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0252] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0253] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0254] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0255] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0256] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0257] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0258] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0259] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0260] This invention is a system for automatically generating data processing workflows based on natural language input from users. The system operates as follows:

[0261] First, the terminal allows users to input information in natural language through its user interface. This natural language input includes information on how the data should be processed and the required output format.

[0262] The server receives this natural language input and analyzes it using a natural language processing engine. Through this analysis, the basic information necessary for the task is extracted. For example, if a user inputs "I want to aggregate sales data by month and create a bar graph," the server identifies two main tasks: "aggregating sales data" and "creating a bar graph."

[0263] Next, the server determines whether further information is needed based on the information obtained. If necessary, the server will present specific questions to the user on the terminal. For example, "What format is the data being handled?" or "Which columns should be included in the aggregation?"

[0264] The user answers these questions and provides additional information. This process allows the server to gain a detailed understanding of the overall workflow.

[0265] The server then references a template library, selects and combines the most suitable workflow templates based on the collected information, and generates a specific workflow. This generated workflow includes steps such as data import, preprocessing, analysis, and visualization.

[0266] Finally, the server executes this workflow. The results of the workflow execution are fed back to the user in real time via the terminal. The feedback includes a success message for the execution, generated graphs, and processed data.

[0267] This system allows users to automatically process data to suit their purposes without relying on specialized tool knowledge. For example, if a marketing researcher wants to quickly grasp trends in sales data, they can simply enter their objective in natural language, and the system will automatically aggregate the data and provide a visually analyzable format.

[0268] The following describes the processing flow.

[0269] Step 1:

[0270] The device displays an interface that prompts the user for natural language input, allowing the user to input tasks and objectives related to data processing in natural language.

[0271] Step 2:

[0272] Users freely input information about their desired tasks and processes using natural language. This includes specifying the type and purpose of the data, as well as the output format.

[0273] Step 3:

[0274] The terminal receives user input and sends that data to the server.

[0275] Step 4:

[0276] The server analyzes the received natural language input using a natural language processing engine to identify basic information and key keywords necessary to accomplish the task.

[0277] Step 5:

[0278] The server identifies missing information from the analysis results and generates questions to supplement the details.

[0279] Step 6:

[0280] The server sends the generated question to the user via the terminal. The question includes data format, processing conditions, and any further details that need to be confirmed.

[0281] Step 7:

[0282] The user enters any additional information required in response to the server's questions and sends it through their terminal.

[0283] Step 8:

[0284] The server analyzes the additional information received from the user and integrates the information so that all the information necessary to generate a complete workflow is available.

[0285] Step 9:

[0286] The server refers to the template library and selects an appropriate workflow template in comparison with the collected information. Then, if necessary, the templates are combined to construct a specific workflow.

[0287] Step 10:

[0288] The server executes the generated workflow and verifies that the process is carried out normally. If an error occurs during the process, the server automatically identifies the problem area and attempts to correct it.

[0289] Step 11:

[0290] The server sends the execution result to the terminal to provide feedback to the user. The feedback includes a success message, generated graphs, processed data, and so on.

[0291] Step 12:

[0292] [[ID=??]] The user checks the result sent through the terminal and performs further analysis and processing if necessary.

[0293] Through this series of steps, the user can effectively and efficiently complete data-related tasks without specialized knowledge of data processing.

[0294] (Example 1)

[0295] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0296] Note: There seems to be a typo in the original text where line ID 3 has "??" instead of a proper ID. It's left as-is in the translation.Data processing tasks often require specialized knowledge and skills, posing a time-consuming and cumbersome problem for non-expert users. Therefore, there is a need for systems that allow users to easily perform data processing and obtain results using natural language requests.

[0297] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0298] In this invention, the server includes a device for receiving input in natural language, a device for analyzing the input natural language and identifying information processing procedures, and a device for querying the user for additional information based on the analysis results. This allows the user to perform data processing in natural language and obtain results efficiently without requiring specialized knowledge.

[0299] A "device that accepts natural language input" is a device equipped with an interface that allows users to input data processing requests into the system using the language they use on a daily basis.

[0300] A "device that analyzes input natural language and identifies information processing procedures" is a device that uses natural language processing technology to extract the tasks and steps necessary for processing from the requests entered by the user.

[0301] A "device that prompts the user for additional information based on analysis results" is a device that interactively asks the user questions when the detailed information necessary for the task identified through analysis is insufficient.

[0302] A "device that selects and combines processing procedures to generate output based on additional information obtained from the user" is a device that utilizes the user's response to select the optimal processing procedure from a template library and combines them to construct a series of data processing tasks.

[0303] The "device that executes the generated processing procedure and provides the execution result to the user via a display device" is a device that executes the assembled processing procedure and provides the result to the user in a form of visualization or display.

[0304] The "generative AI model" is a type of artificial intelligence with functions of natural language analysis and information generation, and is a technology that enables advanced language understanding and response generation.

[0305] The "interactive user interface" is an interface for the user and the system to exchange information with each other, ask appropriate questions or give instructions, and complete tasks sequentially.

[0306] This invention is an information processing system that enables users to perform data processing in natural language without requiring specialized knowledge. In the implementation of this system, analysis using natural language processing technology and an interactive interface play important roles.

[0307] User: The user accesses the system through a terminal and inputs a request related to data processing in natural language. For example, a specific instruction such as "I want to aggregate the sales data monthly and create a bar graph" can be input.

[0308] It is the equipment that receives the natural language input from the user and provides a user interface that can communicate with the database. As a result, the user can confirm the content of their input and get ready to send it to the server. On the terminal, additional input from the user can also be flexibly accepted through the interactive interface. [[ID=]]

[0309] Server: The server uses the natural language input received from the terminal, utilizes a generative AI model (e.g., GPT-3, BERT), and analyzes the input data. Through the analysis, task information required for data processing is extracted. Furthermore, the server refers to the template library, selects an appropriate processing procedure, and automatically generates an optimal workflow based on the user's request.

[0310] Specific example of operation: For instance, a marketing person inputs into the system, "I want to visualize this month's sales data in a bar graph." Based on this request, the server analyzes and applies the sales data processing procedure, and as a result generates the visualized graph.

[0311] By utilizing a generative AI model and prompt messages, processing results are generated in real time and feedback is provided to the user via the device. This system allows users to experience the automation and efficiency improvements of data processing without a steep learning curve.

[0312] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0313] Step 1:

[0314] User: The user inputs data processing requests into the system via a terminal using natural language. This input includes specifying how the data should be aggregated and the output format. An example of input is the text, "I want to aggregate sales data by month and create a bar graph."

[0315] Step 2:

[0316] Terminal: The terminal receives natural language input from the user and sends it to the server. It provides a UI to prompt the user to confirm the input, and after the user confirms the input, it prepares to send it to the server. The input here is the natural language request from step 1, and the output is the data sent to the server.

[0317] Step 3:

[0318] Server: The server analyzes input data from the terminal using a natural language processing engine. It utilizes a generative AI model to decompose and analyze the input text into tokens to identify data processing tasks. The input is natural language text from the terminal, and the output is an internal data structure that defines the task.

[0319] Step 4:

[0320] Server: After the task is identified, the server determines the missing information and generates additional questions as needed. These questions are designed to gather detailed information about the target data format and the data to be aggregated. The input is the analysis result, and the output is the additional question text.

[0321] Step 5:

[0322] Terminal: Displays additional questions from the server to the user and provides an interactive interface for the user to answer. Input is the questions from the server, and output is the user's answer.

[0323] Step 6:

[0324] User: The user enters specific answers to additional questions via their terminal. For example, they might answer, "I want to use data in CSV format and aggregate the year and month columns." The input is the server's question, and the output is the user's additional information.

[0325] Step 7:

[0326] Server: Receives additional user information, references a template library to select an appropriate workflow, and generates specific processing steps. Input is the user's additional information, and output is the automatically generated workflow.

[0327] Step 8:

[0328] Server: Executes the generated workflow and compiles the processing results. Specific operations include data import, preprocessing, analysis, and visualization. The input is the workflow, and the output is the processing result.

[0329] Step 9:

[0330] Terminal: Receives processing results from the server and provides feedback to the user. This includes visualized results such as bar graphs. Input is the processing result from the server, and output is the feedback displayed to the user.

[0331] This process allows users to automatically process data based on simple natural language input and obtain results tailored to their specific needs.

[0332] (Application Example 1)

[0333] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0334] Efficiently processing large amounts of data in cities and providing real-time information useful to citizens and the administration is crucial for optimizing urban management. However, conventional data processing systems require specialized knowledge and make real-time information acquisition and visualization difficult. Furthermore, advanced technical skills were necessary for users to perform data analysis based on their individual needs.

[0335] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0336] In this invention, the server includes means for receiving input in natural language, means for analyzing information from the input natural language to identify basic information, and means for analyzing data acquired in real time and utilizing it as optimized information for visualization in urban activities and business management. As a result, users can perform real-time data processing and analysis of urban data and obtain optimized information by giving instructions in natural language without requiring specialized knowledge.

[0337] "Natural language input" refers to instructions or requests that users make to a computer system using the language that humans use in everyday life.

[0338] "Analysis" is the process of breaking down input natural language data, understanding the information, and extracting the necessary data.

[0339] "Basic information" refers to the data obtained through analysis and the key data and conditions necessary to execute the instructions.

[0340] A "template collection" is a library containing templates for various processing procedures and workflows, which are selected according to specific needs.

[0341] A "work procedure" is a series of processing steps necessary to achieve a specific goal.

[0342] "Real-time feedback" refers to a state where the results of a process are immediately returned to the user, allowing for immediate confirmation and correction.

[0343] "Using it for civic activities and business management" means applying the analyzed data and information in a way that is useful for citizens' lives and administrative operations.

[0344] "Generative artificial intelligence" is a form of artificial intelligence technology that analyzes human natural language and generates flows for instructions and data processing.

[0345] An "interactive interface" refers to a screen or method that allows users to interact with a system and supplement information.

[0346] "Optimizing urban operations" refers to initiatives aimed at efficiently and effectively managing and operating resources and activities within a city.

[0347] To implement this invention, a data processing system based on natural language input is required. The main components of the system include a terminal for receiving natural language input from the user and a server that analyzes the input data and generates the optimal work procedure.

[0348] First, the terminal accepts natural language input from the user via voice or text. This process is carried out via a smartphone or smart glasses, providing an environment where users can easily input instructions. Next, the server analyzes the input using natural language processing technologies such as the Google Cloud Natural Language API and extracts the necessary information for data processing. This analysis clarifies specific tasks, such as updating traffic data or visualizing heatmaps.

[0349] Once the analysis is complete, the server selects the optimal procedure from a set of templates and performs the data processing necessary for city operation and management in real time. The generated results are returned to the terminal in an easy-to-understand format using data visualization libraries such as D3.js and Plotly. This allows users to quickly review the information and take appropriate action.

[0350] As a concrete example, if a user prompts the system with a message such as, "Analyze the air quality data in the city and identify areas that need improvement," the system will immediately acquire the data and generate information useful for urban management. In this way, the invention contributes to optimizing urban operations.

[0351] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0352] Step 1:

[0353] Users input instructions in natural language. For example, a user might use a smartphone or smart glasses to say, "Update the traffic data and display congested areas on a heatmap." At this stage, natural language voice data or text data is obtained as input.

[0354] Step 2:

[0355] The terminal converts received natural language input into text data and sends it to the server. In the case of voice input, speech recognition technology is used to convert it into text. The output here is an instruction in a parseable text format.

[0356] Step 3:

[0357] The server uses the Google Cloud Natural Language API to analyze this text data and extract the necessary tasks and data from the instructions. This process identifies the analyzed instructions and clarifies tasks such as "update traffic data" and "display heatmaps of congested areas." This is the output from the server.

[0358] Step 4:

[0359] Based on the analysis results, the server selects the optimal work procedure from a set of templates and constructs it. At this stage, predefined processing steps within the template library are referenced to select the necessary data acquisition methods and processing steps. The output is the generated workflow.

[0360] Step 5:

[0361] The server retrieves and processes data according to the created workflow. For example, it retrieves the latest traffic data from an API as needed and analyzes it in the specified format. At this point, calculations are performed based on the input data, and data suitable for visualization is generated as output.

[0362] Step 6:

[0363] The server uses visualization libraries such as D3.js and Plotly to visualize the generated data in a user-friendly format. The created heatmaps and analysis results are then fed back to the user's device.

[0364] Step 7:

[0365] The terminal displays real-time feedback from the server to the user. Traffic data is updated as instructed by the user and visually displayed as a heatmap. This allows the user to quickly obtain the necessary information and utilize it.

[0366] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0367] This invention combines a system that automatically generates data processing workflows using natural language input from users with an emotion engine that recognizes user emotions. This system is implemented in the following way.

[0368] First, the device provides a platform for the user to input natural language information through its user interface. This input includes specific purposes and desires regarding data manipulation.

[0369] The server receives natural language input from the terminal and analyzes it using a natural language processing engine. Basic information regarding data processing is extracted here. Simultaneously, the emotion engine is activated to recognize underlying emotions from the user's input. This emotion information is used to adjust the system's response and processing methods.

[0370] Next, the server determines, based on the analysis results, whether any information is missing. If necessary, it also takes into account the emotional state and generates questions to ask the user for additional information in an appropriate tone and expression. These questions are then displayed to the user again through the terminal.

[0371] The user answers questions displayed on the device and provides any necessary additional information. The user's responses are also monitored by an emotion engine. The server integrates this information and generates an appropriate workflow by selecting or combining elements from a template library.

[0372] The generated workflow is executed by the server, and the user's emotions during the process are fed back into the system, dynamically adjusting the workflow. The final processing result is notified to the user via the terminal. The method of presenting this result is also customized according to the user's emotions.

[0373] For example, when a user analyzes sales data, if the system does not meet the user's expectations, it senses this emotion and provides more satisfying results by suggesting different data perspectives or adding process visualizations. Furthermore, when making improvement suggestions, the system takes the user's stress level into consideration and adjusts the information to be presented in a way that is easy for them to accept.

[0374] In this way, we can provide an interactive data processing environment that is adapted to the user's needs and emotions.

[0375] The following describes the processing flow.

[0376] Step 1:

[0377] The terminal provides users with an interface that accepts natural language input, allowing them to freely input their data processing objectives and preferences.

[0378] Step 2:

[0379] The user inputs and sends data processing instructions in natural language via their device, for example, "I want to create a monthly report and analyze trends."

[0380] Step 3:

[0381] The server uses a natural language processing engine to analyze the natural language input it receives and extracts basic information to construct the data processing task.

[0382] Step 4:

[0383] The server simultaneously uses an emotion engine to analyze the user's emotional state based on natural language input and recognize underlying emotions.

[0384] Step 5:

[0385] Based on the basic information and emotional state identified from the analysis results, the server identifies missing information and generates questions to query the user for that information. The tone and content of the questions are adjusted according to the user's emotional state.

[0386] Step 6:

[0387] The server sends the generated question to the terminal and requests additional information from the user.

[0388] Step 7:

[0389] The user enters any necessary additional information in response to the questions displayed on the device and sends it to the server via the device.

[0390] Step 8:

[0391] The server receives additional information provided by the user, refers to a template library to select an appropriate data processing workflow, and then combines and generates it.

[0392] Step 9:

[0393] The server executes the generated workflow while taking the user's emotional state into consideration, and monitors the process to ensure that no problems occur as it progresses.

[0394] Step 10:

[0395] The server collects the execution results, adjusts them according to the user's emotions, and sends them to the terminal for feedback.

[0396] Step 11:

[0397] The device displays feedback sent from the server to the user, presenting it in a format that makes it easy for the user to understand the results.

[0398] This series of steps, through a process that takes user emotions into account, provides more appropriate and flexible data processing results.

[0399] (Example 2)

[0400] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0401] Current data processing systems can adequately analyze users' natural language input to generate optimal data processing workflows, but they struggle to respond to changes in user emotions and incorporate that feedback. As a result, users may have a less satisfying experience, and the system fails to meet expectations for interactive data processing.

[0402] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0403] In this invention, the server includes means for analyzing natural language input, means for recognizing user emotions, means for adjusting responses based on emotions, and means for selecting and generating workflows from a template library. This enables real-time responses to both user input and emotions, and provides optimized data processing workflows.

[0404] "Means of accepting input in natural language" refers to an interface that allows users to provide instructions and information to a system using natural language.

[0405] "Means of analyzing information and identifying basic information related to the subject" refers to the process of analyzing natural language input provided by the user and extracting basic information necessary for data processing from it.

[0406] "Means of querying users for missing information" refers to a mechanism that requests the user, in the form of questions, to complete the missing information identified as a result of the analysis.

[0407] "Means of understanding user emotions" refers to technologies that recognize emotional states from users' natural language input and responses.

[0408] "Means of adjusting response methods based on emotional information" refers to a process that modifies the system's responses and presentation methods according to the recognized emotions of the user, in order to provide an appropriate experience for the user.

[0409] "A means of selecting and combining appropriate workflows from a template library" refers to a method of constructing specific processing procedures by selecting the template that best suits the user's requirements from existing templates, or by combining multiple templates.

[0410] "A means of executing the generated workflow and providing feedback to the user" refers to the process of actually running the constructed data processing workflow and notifying the user of the processing results.

[0411] The system for carrying out this invention operates on a computer network including servers, terminals, and a user interface. Some processing utilizes a cloud-based platform.

[0412] The device provides a user interface, creating a space for users to input information using natural language. This interface allows input via text boxes or voice input, and utilizes tools such as the Google Cloud Natural Language API for natural language processing.

[0413] The server receives natural language input data sent from the terminal and analyzes the information via a natural language processing engine. This engine meticulously analyzes the input sentences and accurately extracts the basic information necessary for data processing. Simultaneously, the server uses sentiment recognition technology to understand the user's emotional state in real time from their input. Sentiment analysis utilizes technologies such as Microsoft Azure Text Analytics.

[0414] Based on emotional information, the server generates questions in an appropriate tone and content to ask the user for missing information, and presents them to the user through the terminal. By utilizing a generative AI model to adjust the prompt text as needed, the quality of communication is improved.

[0415] The user can respond to questions presented on the device and provide any necessary additional information. The emotion recognition engine also operates during this process, providing appropriate feedback based on the user's responses.

[0416] The server integrates all information obtained from the user and selects or generates the most appropriate workflow from the template library. This workflow is dynamically adjusted to match the user's needs based on the generated AI model.

[0417] The processing results obtained from the executed workflow are customized to the user's emotions and ultimately fed back to the user via the device. The results may also be displayed as infographics, providing a visually easy-to-understand format.

[0418] For example, if a user requests, "Find outliers in recent sales data," the system will analyze the input data and ask additional questions about the required data range and calculation criteria. Based on the user's response, it will generate and execute an appropriate workflow and provide feedback on the results in an appropriate format.

[0419] A concrete example of a prompt statement is, "Please tell me the data range required for this process." This allows for the provision of information that aligns with the user's needs and emotions.

[0420] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0421] Step 1:

[0422] The device provides a user interface, allowing users to input information and instructions in natural language. Users can use either a text box or voice input. The input data is captured by the device as natural language text.

[0423] Step 2:

[0424] The server receives natural language text from the terminal as input and sends it to the natural language processing engine. The natural language processing engine analyzes the text and extracts the basic information necessary for data processing. This process includes identifying key keywords and context.

[0425] Step 3:

[0426] Based on the analyzed information, the server uses an emotion recognition engine to recognize the user's emotional state. This input includes natural language text, and the output is either an emotion score or an emotion category. This allows for an understanding of the user's emotions.

[0427] Step 4:

[0428] The server identifies the missing information and generates a question in an appropriate tone, taking sentiment scores into consideration. Using a generative AI model, it creates a prompt and presents a specific question (e.g., "Please tell me the data range required for this process"). This question is then sent to the user.

[0429] Step 5:

[0430] The user enters their answers to the presented questions into the terminal. The additional information entered is specific response data to the questions. This allows for data completion.

[0431] Step 6:

[0432] The server retrieves additional information from the user and selects the most suitable workflow from the template library. Based on the retrieved additional information, data calculations and processing are performed, and the workflow is generated.

[0433] Step 7:

[0434] The server executes the generated workflow. This process performs each step for data processing, and in the process, user sentiment data is re-evaluated. The execution results are then obtained.

[0435] Step 8:

[0436] The server sends the final processing results to the terminal and notifies the user. The results are presented in a customized format based on the user's emotions. This output may take the form of an infographic or summary.

[0437] (Application Example 2)

[0438] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0439] Conventional natural language processing systems do not consider user emotions when analyzing input, and therefore do not optimize responses based on emotions. This can lead to decreased user satisfaction. Furthermore, fixed dialogue methods lack flexibility in sequential information supplementation, making it difficult to provide dynamic information that meets user needs.

[0440] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0441] In this invention, the server includes means for receiving input in natural language, means for analyzing information from the input natural language and identifying basic information related to the subject, and means for recognizing the user's emotions and adjusting the response content according to those emotions. This enables responses that are adapted to the user's emotions, improving the quality of interaction.

[0442] A "means of accepting natural language input" refers to an interface for receiving data and instructions from users in the form of natural language.

[0443] "Means for analyzing information and identifying basic information related to a subject" refers to the process of analyzing the content of natural language received from a user and extracting key points and basic information related to a specific subject based on that content.

[0444] "Means of querying users for missing information" refers to methods of identifying shortcomings in user instructions based on analysis results and asking users about them.

[0445] "Means of recognizing user emotions and adjusting responses accordingly" refers to a mechanism that analyzes the user's emotional state and dynamically changes the content of responses and information presented based on the results.

[0446] "A means of selecting and combining appropriate work procedures from a template library" refers to a system for selecting the most suitable predefined work procedures according to the situation and combining them as needed to construct new procedures.

[0447] "A means of executing the generated work procedure and providing feedback to the user on the results" refers to a function that actually carries out the created work procedure and reports the results and progress of that work to the user.

[0448] This invention relates to a system applied to a consumer robot that provides household support by utilizing natural language input from the user. The system components include a natural language processing engine (NLP engine), an emotion recognition engine, and a task scheduler. The hardware is a home robot equipped with a microphone for receiving voice input, a speaker for outputting responses, and a display for providing visual information.

[0449] The server first receives natural language input from the user through the robot's microphone. The received audio data is analyzed using an NLP engine to understand the user's request. At the same time, the emotion recognition engine is activated to recognize the user's emotional state in real time. This information is used to adjust the response in subsequent processes.

[0450] Once the analysis is complete, the server selects the optimal household chore procedure from the template library and, if necessary, combines and adjusts the order of tasks in a way that adapts to the user's emotions. The created procedure is then executed by a robot, and the progress and results are fed back to the user. The feedback re-evaluates the user's emotions and is reported in an appropriate format that reflects those emotions.

[0451] As a concrete example, suppose a user instructs a robot to "clean the kitchen next Friday." The NLP engine understands this instruction, and the emotion recognition engine detects that the user is tired. In this case, the robot proposes a more efficient cleaning plan and schedules the task in a way that reduces the user's burden.

[0452] Examples of prompts include, "How would the robot respond if the user said, 'When I've finished cleaning the kitchen, suggest a new recipe?'" and "How would you improve the suggestion if the user seemed dissatisfied?" In this way, the system enables efficient and adaptive responses based on the user's instructions and emotions.

[0453] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0454] Step 1:

[0455] The user inputs natural language instructions to the robot via a microphone. This input is voice data, which the terminal collects. This prepares the voice data for transmission to the server.

[0456] Step 2:

[0457] The server provides the audio data received from the terminal to the natural language processing engine, where it is converted into text-based instructions. This text data includes information about specific household tasks and desired actions. This conversion enables analysis in the next step.

[0458] Step 3:

[0459] The server uses the analysis capabilities of its natural language processing engine to extract important information from the text data. It identifies keywords and tasks related to the subject, and then uses an emotion recognition engine to evaluate the user's emotional state. At this stage, input data (text) and intermediate data (emotional information) are generated.

[0460] Step 4:

[0461] The server generates questions to prompt the user for additional information if any is missing. It adjusts the questions, including the appropriate tone and phrasing, taking sentiment recognition results into account. These questions are then sent to the terminal and displayed to the user, setting the stage for obtaining the user's response in the next step.

[0462] Step 5:

[0463] The user provides additional information in response to questions displayed on the device. This response is again collected as audio data and sent to the server. The user's emotional state is also monitored simultaneously.

[0464] Step 6:

[0465] The server integrates additional information and sentiment data received from the user, selects the optimal work procedure from the template library, and combines it as needed to generate a new workflow. This generated workflow is then ready to be executed by the task scheduler.

[0466] Step 7:

[0467] A workflow built by the server is executed by a robot. During the process, the user's emotions are continuously evaluated, and the workflow is dynamically adjusted based on that feedback. Finally, the results are reported to the user via a terminal. This report is presented in a format that reflects the user's emotions.

[0468] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0469] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0470] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0471] [Third Embodiment]

[0472] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0473] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0474] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0475] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0476] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0477] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0478] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0479] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0480] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0481] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0482] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0483] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0484] This invention is a system for automatically generating data processing workflows based on natural language input from users. The system operates as follows:

[0485] First, the terminal allows users to input information in natural language through its user interface. This natural language input includes information on how the data should be processed and the required output format.

[0486] The server receives this natural language input and analyzes it using a natural language processing engine. Through this analysis, the basic information necessary for the task is extracted. For example, if a user inputs "I want to aggregate sales data by month and create a bar graph," the server identifies two main tasks: "aggregating sales data" and "creating a bar graph."

[0487] Next, the server determines whether further information is needed based on the information obtained. If necessary, the server will present specific questions to the user on the terminal. For example, "What format is the data being handled?" or "Which columns should be included in the aggregation?"

[0488] The user answers these questions and provides additional information. This process allows the server to gain a detailed understanding of the overall workflow.

[0489] The server then references a template library, selects and combines the most suitable workflow templates based on the collected information, and generates a specific workflow. This generated workflow includes steps such as data import, preprocessing, analysis, and visualization.

[0490] Finally, the server executes this workflow. The results of the workflow execution are fed back to the user in real time via the terminal. The feedback includes a success message for the execution, generated graphs, and processed data.

[0491] This system allows users to automatically process data to suit their purposes without relying on specialized tool knowledge. For example, if a marketing researcher wants to quickly grasp trends in sales data, they can simply enter their objective in natural language, and the system will automatically aggregate the data and provide a visually analyzable format.

[0492] The following describes the processing flow.

[0493] Step 1:

[0494] The device displays an interface that prompts the user for natural language input, allowing the user to input tasks and objectives related to data processing in natural language.

[0495] Step 2:

[0496] Users freely input information about their desired tasks and processes using natural language. This includes specifying the type and purpose of the data, as well as the output format.

[0497] Step 3:

[0498] The terminal receives user input and sends that data to the server.

[0499] Step 4:

[0500] The server analyzes the received natural language input using a natural language processing engine to identify basic information and key keywords necessary to accomplish the task.

[0501] Step 5:

[0502] The server identifies missing information from the analysis results and generates questions to supplement the details.

[0503] Step 6:

[0504] The server sends the generated question to the user via the terminal. The question includes data format, processing conditions, and any further details that need to be confirmed.

[0505] Step 7:

[0506] The user enters any additional information required in response to the server's questions and sends it through their terminal.

[0507] Step 8:

[0508] The server analyzes the additional information received from the user and integrates the information to ensure that all the necessary information is available to generate a complete workflow.

[0509] Step 9:

[0510] The server refers to the template library, compares it with the collected information, and selects the appropriate workflow template. Then, it combines templates as needed to build the specific workflow.

[0511] Step 10:

[0512] The server executes the generated workflow and verifies that the process runs successfully. If an error occurs during processing, it automatically identifies the problem and attempts to correct it.

[0513] Step 11:

[0514] The server sends the execution results to the terminal and provides feedback to the user. This feedback includes success messages, generated graphs, and processed data.

[0515] Step 12:

[0516] The user reviews the results sent through the device and performs further analysis and processing as needed.

[0517] This series of steps allows users to effectively and efficiently complete data-related tasks, even without specialized data processing knowledge.

[0518] (Example 1)

[0519] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0520] Data processing tasks often require specialized knowledge and skills, posing a time-consuming and cumbersome problem for non-expert users. Therefore, there is a need for systems that allow users to easily perform data processing and obtain results using natural language requests.

[0521] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0522] In this invention, the server includes a device for receiving input in natural language, a device for analyzing the input natural language and identifying information processing procedures, and a device for querying the user for additional information based on the analysis results. This allows the user to perform data processing in natural language and obtain results efficiently without requiring specialized knowledge.

[0523] A "device that accepts natural language input" is a device equipped with an interface that allows users to input data processing requests into the system using the language they use on a daily basis.

[0524] A "device that analyzes input natural language and identifies information processing procedures" is a device that uses natural language processing technology to extract the tasks and steps necessary for processing from the requests entered by the user.

[0525] A "device that prompts the user for additional information based on analysis results" is a device that interactively asks the user questions when the detailed information necessary for the task identified through analysis is insufficient.

[0526] A "device that selects and combines processing procedures to generate output based on additional information obtained from the user" is a device that utilizes the user's response to select the optimal processing procedure from a template library and combines them to construct a series of data processing tasks.

[0527] A "device that executes a generated processing procedure and provides the execution result to the user via a display device" is a device that executes an assembled processing procedure and provides the result to the user in a form that is visualized or displayed.

[0528] A "generative artificial intelligence model" is a type of artificial intelligence that has the function of analyzing natural language and generating information, and is a technology that enables advanced language understanding and response generation.

[0529] An "interactive user interface" is an interface that allows the user and the system to exchange information, ask appropriate questions and give instructions, and complete tasks sequentially.

[0530] This invention is an information processing system that enables users to perform data processing using natural language without requiring specialized knowledge. In implementing this system, analysis using natural language processing technology and an interactive interface play important roles.

[0531] User: Users access the system through a terminal and input data processing requests in natural language. For example, they can input a specific command such as, "I want to aggregate sales data by month and create a bar graph."

[0532] Terminal: The terminal receives natural language input from the user and provides a user interface that can communicate with the database. This allows the user to review their input and prepare it for transmission to the server. The interactive interface on the terminal can also flexibly accept additional user input.

[0533] Server: The server uses natural language input received from the terminal to analyze the input data using generative AI models (e.g., GPT-3, BERT). Through this analysis, task information necessary for data processing is extracted. Furthermore, the server refers to a template library to select appropriate processing steps and automatically generates an optimal workflow based on the user's requirements.

[0534] Specific example of operation: For instance, a marketing person inputs into the system, "I want to visualize this month's sales data in a bar graph." Based on this request, the server analyzes and applies the sales data processing procedure, and as a result generates the visualized graph.

[0535] By utilizing a generative AI model and prompt messages, processing results are generated in real time and feedback is provided to the user via the terminal. This system allows users to experience the automation and efficiency improvements of data processing without a steep learning curve.

[0536] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0537] Step 1:

[0538] User: The user inputs data processing requests into the system via a terminal using natural language. This input includes specifying how the data should be aggregated and the output format. An example of input is the text, "I want to aggregate sales data by month and create a bar graph."

[0539] Step 2:

[0540] Terminal: The terminal receives natural language input from the user and sends it to the server. It provides a UI to prompt the user to confirm the input, and after the user confirms the input, it prepares to send it to the server. The input here is the natural language request from step 1, and the output is the data sent to the server.

[0541] Step 3:

[0542] Server: The server analyzes input data from the terminal using a natural language processing engine. It utilizes a generative AI model to decompose and analyze the input text into tokens to identify data processing tasks. The input is natural language text from the terminal, and the output is an internal data structure that defines the task.

[0543] Step 4:

[0544] Server: After the task is identified, the server determines the missing information and generates additional questions as needed. These questions are designed to gather detailed information about the target data format and the data to be aggregated. The input is the analysis result, and the output is the additional question text.

[0545] Step 5:

[0546] Terminal: Displays additional questions from the server to the user and provides an interactive interface for the user to answer. Input is the questions from the server, and output is the user's answer.

[0547] Step 6:

[0548] User: The user enters specific answers to additional questions via their terminal. For example, they might answer, "I want to use data in CSV format and aggregate the year and month columns." The input is the server's question, and the output is the user's additional information.

[0549] Step 7:

[0550] Server: Receives additional user information, references a template library to select an appropriate workflow, and generates specific processing steps. Input is the user's additional information, and output is the automatically generated workflow.

[0551] Step 8:

[0552] Server: Executes the generated workflow and compiles the processing results. Specific operations include data import, preprocessing, analysis, and visualization. The input is the workflow, and the output is the processing result.

[0553] Step 9:

[0554] Terminal: Receives processing results from the server and provides feedback to the user. This includes visualized results such as bar graphs. Input is the processing result from the server, and output is the feedback displayed to the user.

[0555] This process allows users to automatically process data based on simple natural language input and obtain results tailored to their specific needs.

[0556] (Application Example 1)

[0557] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0558] Efficiently processing large amounts of data in cities and providing real-time information useful to citizens and the administration is crucial for optimizing urban management. However, conventional data processing systems require specialized knowledge and make real-time information acquisition and visualization difficult. Furthermore, advanced technical skills were necessary for users to perform data analysis based on their individual needs.

[0559] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0560] In this invention, the server includes means for receiving input in natural language, means for analyzing information from the input natural language to identify basic information, and means for analyzing data acquired in real time and utilizing it as optimized information for visualization in urban activities and business management. As a result, users can perform real-time data processing and analysis of urban data and obtain optimized information by giving instructions in natural language without requiring specialized knowledge.

[0561] "Natural language input" refers to instructions or requests that users make to a computer system using the language that humans use in everyday life.

[0562] "Analysis" is the process of breaking down input natural language data, understanding the information, and extracting the necessary data.

[0563] "Basic information" refers to the data obtained through analysis and the key data and conditions necessary to execute the instructions.

[0564] A "template collection" is a library containing templates for various processing procedures and workflows, which are selected according to specific needs.

[0565] A "work procedure" is a series of processing steps necessary to achieve a specific goal.

[0566] "Real-time feedback" refers to a state where the results of a process are immediately returned to the user, allowing for immediate confirmation and correction.

[0567] "Using it for civic activities and business management" means applying the analyzed data and information in a way that is useful for citizens' lives and administrative operations.

[0568] "Generative artificial intelligence" is a form of artificial intelligence technology that analyzes human natural language and generates flows for instructions and data processing.

[0569] An "interactive interface" refers to a screen or method that allows users to interact with a system and supplement information.

[0570] "Optimizing urban operations" refers to initiatives aimed at efficiently and effectively managing and operating resources and activities within a city.

[0571] To implement this invention, a data processing system based on natural language input is required. The main components of the system include a terminal for receiving natural language input from the user and a server that analyzes the input data and generates the optimal work procedure.

[0572] First, the terminal accepts natural language input from the user via voice or text. This process is carried out via a smartphone or smart glasses, providing an environment where users can easily input instructions. Next, the server analyzes the input using natural language processing technologies such as the Google Cloud Natural Language API and extracts the necessary information for data processing. This analysis clarifies specific tasks, such as updating traffic data or visualizing heatmaps.

[0573] Once the analysis is complete, the server selects the optimal procedure from a set of templates and performs the data processing necessary for city operation and management in real time. The generated results are returned to the terminal in an easy-to-understand format using data visualization libraries such as D3.js and Plotly. This allows users to quickly review the information and take appropriate action.

[0574] As a concrete example, if a user prompts the system with a message such as, "Analyze the air quality data in the city and identify areas that need improvement," the system will immediately acquire the data and generate information useful for urban management. In this way, the invention contributes to optimizing urban operations.

[0575] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0576] Step 1:

[0577] Users input instructions in natural language. For example, a user might use a smartphone or smart glasses to say, "Update the traffic data and display congested areas on a heatmap." At this stage, natural language voice data or text data is obtained as input.

[0578] Step 2:

[0579] The terminal converts received natural language input into text data and sends it to the server. In the case of voice input, speech recognition technology is used to convert it into text. The output here is an instruction in a parseable text format.

[0580] Step 3:

[0581] The server uses the Google Cloud Natural Language API to analyze this text data and extract the necessary tasks and data from the instructions. This process identifies the analyzed instructions and clarifies tasks such as "update traffic data" and "display heatmaps of congested areas." This is the output from the server.

[0582] Step 4:

[0583] Based on the analysis results, the server selects the optimal work procedure from a set of templates and constructs it. At this stage, predefined processing steps within the template library are referenced to select the necessary data acquisition methods and processing steps. The output is the generated workflow.

[0584] Step 5:

[0585] The server retrieves and processes data according to the created workflow. For example, it retrieves the latest traffic data from an API as needed and analyzes it in the specified format. At this point, calculations are performed based on the input data, and data suitable for visualization is generated as output.

[0586] Step 6:

[0587] The server uses visualization libraries such as D3.js and Plotly to visualize the generated data in a user-friendly format. The created heatmaps and analysis results are then fed back to the user's device.

[0588] Step 7:

[0589] The terminal displays real-time feedback from the server to the user. Traffic data is updated as instructed by the user and visually displayed as a heatmap. This allows the user to quickly obtain the necessary information and utilize it.

[0590] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0591] This invention combines a system that automatically generates data processing workflows using natural language input from users with an emotion engine that recognizes user emotions. This system is implemented in the following way.

[0592] First, the device provides a platform for the user to input natural language through its user interface. This input includes specific purposes and desires regarding data manipulation.

[0593] The server receives natural language input from the terminal and analyzes it using a natural language processing engine. Basic information regarding data processing is extracted here. Simultaneously, the emotion engine is activated to recognize underlying emotions from the user's input. This emotion information is used to adjust the system's response and processing methods.

[0594] Next, the server determines, based on the analysis results, whether any information is missing. If necessary, it also takes into account the emotional state and generates questions to ask the user for additional information in an appropriate tone and expression. These questions are then displayed to the user again through the terminal.

[0595] The user answers questions displayed on the device and provides any necessary additional information. The user's responses are also monitored by an emotion engine. The server integrates this information and generates an appropriate workflow by selecting or combining elements from a template library.

[0596] The generated workflow is executed by the server, and the user's emotions during the process are fed back into the system, dynamically adjusting the workflow. The final processing result is notified to the user via the terminal. The method of presenting this result is also customized according to the user's emotions.

[0597] For example, when a user analyzes sales data, if the system does not meet the user's expectations, it senses this emotion and provides more satisfying results by suggesting different data perspectives or adding process visualizations. Furthermore, when making improvement suggestions, the system takes the user's stress level into consideration and adjusts the information to be presented in a way that is easy for them to accept.

[0598] In this way, we can provide an interactive data processing environment that is adapted to the user's needs and emotions.

[0599] The following describes the processing flow.

[0600] Step 1:

[0601] The terminal provides users with an interface that accepts natural language input, allowing them to freely input their data processing objectives and preferences.

[0602] Step 2:

[0603] The user inputs and sends data processing instructions in natural language via their device, for example, "I want to create a monthly report and analyze trends."

[0604] Step 3:

[0605] The server uses a natural language processing engine to analyze the natural language input it receives and extracts basic information to construct the data processing task.

[0606] Step 4:

[0607] The server simultaneously uses an emotion engine to analyze the user's emotional state based on natural language input and recognize underlying emotions.

[0608] Step 5:

[0609] Based on the basic information and emotional state identified from the analysis results, the server identifies missing information and generates questions to query the user for that information. The tone and content of the questions are adjusted according to the user's emotional state.

[0610] Step 6:

[0611] The server sends the generated question to the terminal and requests additional information from the user.

[0612] Step 7:

[0613] The user enters any necessary additional information in response to the questions displayed on the device and sends it to the server via the device.

[0614] Step 8:

[0615] The server receives additional information provided by the user, refers to a template library to select an appropriate data processing workflow, and then combines and generates it.

[0616] Step 9:

[0617] The server executes the generated workflow while taking into account the user's emotional state, and monitors the process to ensure that no problems occur as it progresses.

[0618] Step 10:

[0619] The server collects the execution results, adjusts them according to the user's emotions, and sends them to the terminal for feedback.

[0620] Step 11:

[0621] The device displays feedback sent from the server to the user, presenting it in a format that makes it easy for the user to understand the results.

[0622] This series of steps, through a process that takes user emotions into account, provides more appropriate and flexible data processing results.

[0623] (Example 2)

[0624] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0625] Current data processing systems can adequately analyze users' natural language input to generate optimal data processing workflows, but they struggle to respond to changes in user emotions and incorporate that feedback. As a result, users may have a less satisfying experience, and the system fails to meet expectations for interactive data processing.

[0626] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0627] In this invention, the server includes means for analyzing natural language input, means for recognizing user emotions, means for adjusting responses based on emotions, and means for selecting and generating workflows from a template library. This enables real-time responses to both user input and emotions, and provides optimized data processing workflows.

[0628] "Means of accepting input in natural language" refers to an interface that allows users to provide instructions and information to a system using natural language.

[0629] "Means of analyzing information and identifying basic information related to the subject" refers to the process of analyzing natural language input provided by the user and extracting basic information necessary for data processing from it.

[0630] "Means of querying users for missing information" refers to a mechanism that requests the user, in the form of questions, to complete the missing information identified as a result of the analysis.

[0631] "Means of understanding user emotions" refers to technologies that recognize emotional states from users' natural language input and responses.

[0632] "Means of adjusting response methods based on emotional information" refers to a process that modifies the system's responses and presentation methods according to the recognized emotions of the user, in order to provide an appropriate experience for the user.

[0633] "A means of selecting and combining appropriate workflows from a template library" refers to a method of constructing specific processing procedures by selecting the template that best suits the user's requirements from existing templates, or by combining multiple templates.

[0634] "A means of executing the generated workflow and providing feedback to the user" refers to the process of actually running the constructed data processing workflow and notifying the user of the processing results.

[0635] The system for carrying out this invention operates on a computer network including servers, terminals, and a user interface. Some processing utilizes a cloud-based platform.

[0636] The device provides a user interface, creating a space for users to input information using natural language. This interface allows input via text boxes or voice input, and utilizes tools such as the Google Cloud Natural Language API for natural language processing.

[0637] The server receives natural language input data sent from the terminal and analyzes the information via a natural language processing engine. This engine meticulously analyzes the input sentences and accurately extracts the basic information necessary for data processing. Simultaneously, the server uses sentiment recognition technology to understand the user's emotional state in real time from their input. Sentiment analysis utilizes technologies such as Microsoft Azure Text Analytics.

[0638] Based on emotional information, the server generates questions in an appropriate tone and content to ask the user for missing information, and presents them to the user through the terminal. By utilizing a generative AI model to adjust the prompt text as needed, the quality of communication is improved.

[0639] The user can respond to questions presented on the device and provide any necessary additional information. The emotion recognition engine also operates during this process, providing appropriate feedback based on the user's responses.

[0640] The server integrates all information obtained from the user and selects or generates the most appropriate workflow from the template library. This workflow is dynamically adjusted to match the user's needs based on the generated AI model.

[0641] The processing results obtained from the executed workflow are customized to the user's emotions and ultimately fed back to the user via the device. The results may also be displayed as infographics, providing a visually easy-to-understand format.

[0642] For example, if a user requests, "Find outliers in recent sales data," the system will analyze the input data and ask additional questions about the required data range and calculation criteria. Based on the user's response, it will generate and execute an appropriate workflow and provide feedback on the results in an appropriate format.

[0643] A concrete example of a prompt statement is, "Please tell me the data range required for this process." This allows for the provision of information that aligns with the user's needs and emotions.

[0644] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0645] Step 1:

[0646] The device provides a user interface, allowing users to input information and instructions in natural language. Users can use either a text box or voice input. The input data is captured by the device as natural language text.

[0647] Step 2:

[0648] The server receives natural language text from the terminal as input and sends it to the natural language processing engine. The natural language processing engine analyzes the text and extracts the basic information necessary for data processing. This process includes identifying key keywords and context.

[0649] Step 3:

[0650] Based on the analyzed information, the server uses an emotion recognition engine to recognize the user's emotional state. This input includes natural language text, and the output is either an emotion score or an emotion category. This allows for an understanding of the user's emotions.

[0651] Step 4:

[0652] The server identifies the missing information and generates a question in an appropriate tone, taking sentiment scores into consideration. Using a generative AI model, it creates a prompt and presents a specific question (e.g., "Please tell me the data range required for this process"). This question is then sent to the user.

[0653] Step 5:

[0654] The user enters their answers to the presented questions into the terminal. The additional information entered is specific response data to the questions. This allows for data completion.

[0655] Step 6:

[0656] The server retrieves additional information from the user and selects the most suitable workflow from the template library. Based on the retrieved additional information, data calculations and processing are performed, and the workflow is generated.

[0657] Step 7:

[0658] The server executes the generated workflow. This process performs each step for data processing, and in the process, user sentiment data is re-evaluated. The execution results are then obtained.

[0659] Step 8:

[0660] The server sends the final processing results to the terminal and notifies the user. The results are presented in a customized format based on the user's emotions. This output may take the form of an infographic or summary.

[0661] (Application Example 2)

[0662] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0663] Conventional natural language processing systems do not consider user emotions when analyzing input, and therefore do not optimize responses based on emotions. This can lead to decreased user satisfaction. Furthermore, fixed dialogue methods lack flexibility in sequential information supplementation, making it difficult to provide dynamic information that meets user needs.

[0664] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0665] In this invention, the server includes means for receiving input in natural language, means for analyzing information from the input natural language and identifying basic information related to the subject, and means for recognizing the user's emotions and adjusting the response content according to those emotions. This enables responses that are adapted to the user's emotions, improving the quality of interaction.

[0666] A "means of accepting natural language input" refers to an interface for receiving data and instructions from users in the form of natural language.

[0667] "Means for analyzing information and identifying basic information related to a subject" refers to the process of analyzing the content of natural language received from a user and extracting key points and basic information related to a specific subject based on that content.

[0668] "Means of querying users for missing information" refers to methods of identifying shortcomings in user instructions based on analysis results and asking users about them.

[0669] "Means of recognizing user emotions and adjusting responses accordingly" refers to a mechanism that analyzes the user's emotional state and dynamically changes the content of responses and information presented based on the results.

[0670] "A means of selecting and combining appropriate work procedures from a template library" refers to a system for selecting the most suitable predefined work procedures according to the situation and combining them as needed to construct new procedures.

[0671] "A means of executing the generated work procedure and providing feedback to the user on the results" refers to a function that actually carries out the created work procedure and reports the results and progress of that work to the user.

[0672] This invention relates to a system applied to a consumer robot that provides household support by utilizing natural language input from the user. The system components include a natural language processing engine (NLP engine), an emotion recognition engine, and a task scheduler. The hardware is a home robot equipped with a microphone for receiving voice input, a speaker for outputting responses, and a display for providing visual information.

[0673] The server first receives natural language input from the user through the robot's microphone. The received audio data is analyzed using an NLP engine to understand the user's request. At the same time, the emotion recognition engine is activated to recognize the user's emotional state in real time. This information is used to adjust the response in subsequent processes.

[0674] Once the analysis is complete, the server selects the optimal household chore procedure from the template library and, if necessary, combines and adjusts the order of tasks in a way that adapts to the user's emotions. The created procedure is then executed by a robot, and the progress and results are fed back to the user. The feedback re-evaluates the user's emotions and is reported in an appropriate format that reflects those emotions.

[0675] As a concrete example, suppose a user instructs a robot to "clean the kitchen next Friday." The NLP engine understands this instruction, and the emotion recognition engine detects that the user is tired. In this case, the robot proposes a more efficient cleaning plan and schedules the task in a way that reduces the user's burden.

[0676] Examples of prompts include, "How would the robot respond if the user said, 'When I've finished cleaning the kitchen, suggest a new recipe?'" and "How would you improve the suggestion if the user seemed dissatisfied?" In this way, the system enables efficient and adaptive responses based on the user's instructions and emotions.

[0677] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0678] Step 1:

[0679] The user inputs natural language instructions to the robot via a microphone. This input is voice data, which the terminal collects. This prepares the voice data for transmission to the server.

[0680] Step 2:

[0681] The server provides the audio data received from the terminal to the natural language processing engine, where it is converted into text-based instructions. This text data includes information about specific household tasks and desired actions. This conversion enables analysis in the next step.

[0682] Step 3:

[0683] The server uses the analysis capabilities of its natural language processing engine to extract important information from the text data. It identifies keywords and tasks related to the subject, and then uses an emotion recognition engine to evaluate the user's emotional state. At this stage, input data (text) and intermediate data (emotional information) are generated.

[0684] Step 4:

[0685] The server generates questions to prompt the user for additional information if any is missing. It adjusts the questions, including the appropriate tone and phrasing, taking sentiment recognition results into account. These questions are then sent to the terminal and displayed to the user, setting the stage for obtaining the user's response in the next step.

[0686] Step 5:

[0687] The user provides additional information in response to questions displayed on the device. This response is again collected as audio data and sent to the server. The user's emotional state is also monitored simultaneously.

[0688] Step 6:

[0689] The server integrates additional information and sentiment data received from the user, selects the optimal work procedure from the template library, and combines it as needed to generate a new workflow. This generated workflow is then ready to be executed by the task scheduler.

[0690] Step 7:

[0691] A workflow built by the server is executed by a robot. During the process, the user's emotions are continuously evaluated, and the workflow is dynamically adjusted based on that feedback. Finally, the results are reported to the user via a terminal. This report is presented in a format that reflects the user's emotions.

[0692] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0693] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0694] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0695] [Fourth Embodiment]

[0696] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0697] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0698] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0699] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0700] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0701] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0702] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0703] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0704] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0705] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0706] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0707] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0708] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0709] This invention is a system for automatically generating data processing workflows based on natural language input from users. The system operates as follows:

[0710] First, the terminal allows users to input information in natural language through its user interface. This natural language input includes information on how the data should be processed and the required output format.

[0711] The server receives this natural language input and analyzes it using a natural language processing engine. Through this analysis, the basic information necessary for the task is extracted. For example, if a user inputs "I want to aggregate sales data by month and create a bar graph," the server identifies two main tasks: "aggregating sales data" and "creating a bar graph."

[0712] Next, the server determines whether further information is needed based on the information obtained. If necessary, the server will present specific questions to the user on the terminal. For example, "What format is the data being handled?" or "Which columns should be included in the aggregation?"

[0713] The user answers these questions and provides additional information. This process allows the server to gain a detailed understanding of the overall workflow.

[0714] The server then references a template library, selects and combines the most suitable workflow templates based on the collected information, and generates a specific workflow. This generated workflow includes steps such as data import, preprocessing, analysis, and visualization.

[0715] Finally, the server executes this workflow. The results of the workflow execution are fed back to the user in real time via the terminal. The feedback includes a success message for the execution, generated graphs, and processed data.

[0716] This system allows users to automatically process data to suit their purposes without relying on specialized tool knowledge. For example, if a marketing researcher wants to quickly grasp trends in sales data, they can simply enter their objective in natural language, and the system will automatically aggregate the data and provide a visually analyzable format.

[0717] The following describes the processing flow.

[0718] Step 1:

[0719] The device displays an interface that prompts the user for natural language input, allowing the user to input tasks and objectives related to data processing in natural language.

[0720] Step 2:

[0721] Users freely input information about their desired tasks and processes using natural language. This includes specifying the type and purpose of the data, as well as the output format.

[0722] Step 3:

[0723] The terminal receives user input and sends that data to the server.

[0724] Step 4:

[0725] The server analyzes the received natural language input using a natural language processing engine to identify basic information and key keywords necessary to accomplish the task.

[0726] Step 5:

[0727] The server identifies missing information from the analysis results and generates questions to supplement the details.

[0728] Step 6:

[0729] The server sends the generated question to the user via the terminal. The question includes data format, processing conditions, and any further details that need to be confirmed.

[0730] Step 7:

[0731] The user enters any additional information required in response to the server's questions and sends it through their terminal.

[0732] Step 8:

[0733] The server analyzes the additional information received from the user and integrates the information to ensure that all the necessary information is available to generate a complete workflow.

[0734] Step 9:

[0735] The server refers to the template library, compares it with the collected information, and selects the appropriate workflow template. Then, it combines templates as needed to build the specific workflow.

[0736] Step 10:

[0737] The server executes the generated workflow and verifies that the process runs successfully. If an error occurs during processing, it automatically identifies the problem and attempts to correct it.

[0738] Step 11:

[0739] The server sends the execution results to the terminal and provides feedback to the user. This feedback includes success messages, generated graphs, and processed data.

[0740] Step 12:

[0741] The user reviews the results sent through the device and performs further analysis and processing as needed.

[0742] This series of steps allows users to effectively and efficiently complete data-related tasks, even without specialized data processing knowledge.

[0743] (Example 1)

[0744] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0745] Data processing tasks often require specialized knowledge and skills, posing a time-consuming and cumbersome problem for non-expert users. Therefore, there is a need for systems that allow users to easily perform data processing and obtain results using natural language requests.

[0746] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0747] In this invention, the server includes a device for receiving input in natural language, a device for analyzing the input natural language and identifying information processing procedures, and a device for querying the user for additional information based on the analysis results. This allows the user to perform data processing in natural language and obtain results efficiently without requiring specialized knowledge.

[0748] A "device that accepts natural language input" is a device equipped with an interface that allows users to input data processing requests into the system using the language they use on a daily basis.

[0749] A "device that analyzes input natural language and identifies information processing procedures" is a device that uses natural language processing technology to extract the tasks and steps necessary for processing from the requests entered by the user.

[0750] A "device that prompts the user for additional information based on analysis results" is a device that interactively asks the user questions when the detailed information necessary for the task identified through analysis is insufficient.

[0751] A "device that selects and combines processing procedures to generate output based on additional information obtained from the user" is a device that utilizes the user's response to select the optimal processing procedure from a template library and combines them to construct a series of data processing tasks.

[0752] A "device that executes a generated processing procedure and provides the execution result to the user via a display device" is a device that executes an assembled processing procedure and provides the result to the user in a form that is visualized or displayed.

[0753] A "generative artificial intelligence model" is a type of artificial intelligence that has the function of analyzing natural language and generating information, and is a technology that enables advanced language understanding and response generation.

[0754] An "interactive user interface" is an interface that allows the user and the system to exchange information, ask appropriate questions and give instructions, and complete tasks sequentially.

[0755] This invention is an information processing system that enables users to perform data processing using natural language without requiring specialized knowledge. In implementing this system, analysis using natural language processing technology and an interactive interface play important roles.

[0756] User: Users access the system through a terminal and input data processing requests in natural language. For example, they can input a specific command such as, "I want to aggregate sales data by month and create a bar graph."

[0757] Terminal: The terminal receives natural language input from the user and provides a user interface that can communicate with the database. This allows the user to review their input and prepare it for transmission to the server. The interactive interface on the terminal can also flexibly accept additional user input.

[0758] Server: The server uses natural language input received from the terminal to analyze the input data using generative AI models (e.g., GPT-3, BERT). Through this analysis, task information necessary for data processing is extracted. Furthermore, the server refers to a template library to select appropriate processing steps and automatically generates an optimal workflow based on the user's requirements.

[0759] Specific example of operation: For instance, a marketing person inputs into the system, "I want to visualize this month's sales data in a bar graph." Based on this request, the server analyzes and applies the sales data processing procedure, and as a result generates the visualized graph.

[0760] By utilizing a generative AI model and prompt messages, processing results are generated in real time and feedback is provided to the user via the terminal. This system allows users to experience the automation and efficiency improvements of data processing without a steep learning curve.

[0761] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0762] Step 1:

[0763] User: The user inputs data processing requests into the system via a terminal using natural language. This input includes specifying how the data should be aggregated and the output format. An example of input is the text, "I want to aggregate sales data by month and create a bar graph."

[0764] Step 2:

[0765] Terminal: The terminal receives natural language input from the user and sends it to the server. It provides a UI to prompt the user to confirm the input, and after the user confirms the input, it prepares to send it to the server. The input here is the natural language request from step 1, and the output is the data sent to the server.

[0766] Step 3:

[0767] Server: The server analyzes input data from the terminal using a natural language processing engine. It utilizes a generative AI model to decompose and analyze the input text into tokens to identify data processing tasks. The input is natural language text from the terminal, and the output is an internal data structure that defines the task.

[0768] Step 4:

[0769] Server: After the task is identified, the server determines the missing information and generates additional questions as needed. These questions are designed to gather detailed information about the target data format and the data to be aggregated. The input is the analysis result, and the output is the additional question text.

[0770] Step 5:

[0771] Terminal: Displays additional questions from the server to the user and provides an interactive interface for the user to answer. Input is the questions from the server, and output is the user's answer.

[0772] Step 6:

[0773] User: The user enters specific answers to additional questions via their terminal. For example, they might answer, "I want to use data in CSV format and aggregate the year and month columns." The input is the server's question, and the output is the user's additional information.

[0774] Step 7:

[0775] Server: Receives additional user information, references a template library to select an appropriate workflow, and generates specific processing steps. Input is the user's additional information, and output is the automatically generated workflow.

[0776] Step 8:

[0777] Server: Executes the generated workflow and compiles the processing results. Specific operations include data import, preprocessing, analysis, and visualization. The input is the workflow, and the output is the processing result.

[0778] Step 9:

[0779] Terminal: Receives processing results from the server and provides feedback to the user. This includes visualized results such as bar graphs. Input is the processing result from the server, and output is the feedback displayed to the user.

[0780] This process allows users to automatically process data based on simple natural language input and obtain results tailored to their specific needs.

[0781] (Application Example 1)

[0782] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0783] Efficiently processing large amounts of data in cities and providing real-time information useful to citizens and the administration is crucial for optimizing urban management. However, conventional data processing systems require specialized knowledge and make real-time information acquisition and visualization difficult. Furthermore, advanced technical skills were necessary for users to perform data analysis based on their individual needs.

[0784] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0785] In this invention, the server includes means for receiving input in natural language, means for analyzing information from the input natural language to identify basic information, and means for analyzing data acquired in real time and utilizing it as optimized information for visualization in urban activities and business management. As a result, users can perform real-time data processing and analysis of urban data and obtain optimized information by giving instructions in natural language without requiring specialized knowledge.

[0786] "Natural language input" refers to instructions or requests that users make to a computer system using the language that humans use in everyday life.

[0787] "Analysis" is the process of breaking down input natural language data, understanding the information, and extracting the necessary data.

[0788] "Basic information" refers to the data obtained through analysis and the key data and conditions necessary to execute the instructions.

[0789] A "template collection" is a library containing templates for various processing procedures and workflows, which are selected according to specific needs.

[0790] A "work procedure" is a series of processing steps necessary to achieve a specific goal.

[0791] "Real-time feedback" refers to a state where the results of a process are immediately returned to the user, allowing for immediate confirmation and correction.

[0792] "Using it for civic activities and business management" means applying the analyzed data and information in a way that is useful for citizens' lives and administrative operations.

[0793] "Generative artificial intelligence" is a form of artificial intelligence technology that analyzes human natural language and generates flows for instructions and data processing.

[0794] An "interactive interface" refers to a screen or method that allows users to interact with a system and supplement information.

[0795] "Optimizing urban operations" refers to initiatives aimed at efficiently and effectively managing and operating resources and activities within a city.

[0796] To implement this invention, a data processing system based on natural language input is required. The main components of the system include a terminal for receiving natural language input from the user and a server that analyzes the input data and generates the optimal work procedure.

[0797] First, the terminal accepts natural language input from the user via voice or text. This process is carried out via a smartphone or smart glasses, providing an environment where users can easily input instructions. Next, the server analyzes the input using natural language processing technologies such as the Google Cloud Natural Language API and extracts the necessary information for data processing. This analysis clarifies specific tasks, such as updating traffic data or visualizing heatmaps.

[0798] Once the analysis is complete, the server selects the optimal procedure from a set of templates and performs the data processing necessary for city operation and management in real time. The generated results are returned to the terminal in an easy-to-understand format using data visualization libraries such as D3.js and Plotly. This allows users to quickly review the information and take appropriate action.

[0799] As a concrete example, if a user prompts the system with a message such as, "Analyze the air quality data in the city and identify areas that need improvement," the system will immediately acquire the data and generate information useful for urban management. In this way, the invention contributes to optimizing urban operations.

[0800] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0801] Step 1:

[0802] Users input instructions in natural language. For example, a user might use a smartphone or smart glasses to say, "Update the traffic data and display congested areas on a heatmap." At this stage, natural language voice data or text data is obtained as input.

[0803] Step 2:

[0804] The terminal converts received natural language input into text data and sends it to the server. In the case of voice input, speech recognition technology is used to convert it into text. The output here is an instruction in a parseable text format.

[0805] Step 3:

[0806] The server uses the Google Cloud Natural Language API to analyze this text data and extract the necessary tasks and data from the instructions. This process identifies the analyzed instructions and clarifies tasks such as "update traffic data" and "display heatmaps of congested areas." This is the output from the server.

[0807] Step 4:

[0808] Based on the analysis results, the server selects the optimal work procedure from a set of templates and constructs it. At this stage, predefined processing steps within the template library are referenced to select the necessary data acquisition methods and processing steps. The output is the generated workflow.

[0809] Step 5:

[0810] The server retrieves and processes data according to the created workflow. For example, it retrieves the latest traffic data from an API as needed and analyzes it in the specified format. At this point, calculations are performed based on the input data, and data suitable for visualization is generated as output.

[0811] Step 6:

[0812] The server uses visualization libraries such as D3.js and Plotly to visualize the generated data in a user-friendly format. The created heatmaps and analysis results are then fed back to the user's device.

[0813] Step 7:

[0814] The terminal displays real-time feedback from the server to the user. Traffic data is updated as instructed by the user and visually displayed as a heatmap. This allows the user to quickly obtain the necessary information and utilize it.

[0815] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0816] This invention combines a system that automatically generates data processing workflows using natural language input from users with an emotion engine that recognizes user emotions. This system is implemented in the following way.

[0817] First, the device provides a platform for the user to input natural language through its user interface. This input includes specific purposes and desires regarding data manipulation.

[0818] The server receives natural language input from the terminal and analyzes it using a natural language processing engine. Basic information regarding data processing is extracted here. Simultaneously, the emotion engine is activated to recognize underlying emotions from the user's input. This emotion information is used to adjust the system's response and processing methods.

[0819] Next, the server determines, based on the analysis results, whether any information is missing. If necessary, it also takes into account the emotional state and generates questions to ask the user for additional information in an appropriate tone and expression. These questions are then displayed to the user again through the terminal.

[0820] The user answers questions displayed on the device and provides any necessary additional information. The user's responses are also monitored by an emotion engine. The server integrates this information and generates an appropriate workflow by selecting or combining elements from a template library.

[0821] The generated workflow is executed by the server, and the user's emotions during the process are fed back into the system, dynamically adjusting the workflow. The final processing result is notified to the user via the terminal. The method of presenting this result is also customized according to the user's emotions.

[0822] For example, when a user analyzes sales data, if the system does not meet the user's expectations, it senses this emotion and provides more satisfying results by suggesting different data perspectives or adding process visualizations. Furthermore, when making improvement suggestions, the system takes the user's stress level into consideration and adjusts the information to be presented in a way that is easy for them to accept.

[0823] In this way, we can provide an interactive data processing environment that is adapted to the user's needs and emotions.

[0824] The following describes the processing flow.

[0825] Step 1:

[0826] The terminal provides users with an interface that accepts natural language input, allowing them to freely input their data processing objectives and preferences.

[0827] Step 2:

[0828] The user inputs and sends data processing instructions in natural language via their device, for example, "I want to create a monthly report and analyze trends."

[0829] Step 3:

[0830] The server uses a natural language processing engine to analyze the natural language input it receives and extracts basic information to construct the data processing task.

[0831] Step 4:

[0832] The server simultaneously uses an emotion engine to analyze the user's emotional state based on natural language input and recognize underlying emotions.

[0833] Step 5:

[0834] Based on the basic information and emotional state identified from the analysis results, the server identifies missing information and generates questions to query the user for that information. The tone and content of the questions are adjusted according to the user's emotional state.

[0835] Step 6:

[0836] The server sends the generated question to the terminal and requests additional information from the user.

[0837] Step 7:

[0838] The user enters any necessary additional information in response to the questions displayed on the device and sends it to the server via the device.

[0839] Step 8:

[0840] The server receives additional information provided by the user, refers to a template library to select an appropriate data processing workflow, and then combines and generates it.

[0841] Step 9:

[0842] The server executes the generated workflow while taking into account the user's emotional state, and monitors the process to ensure that no problems occur as it progresses.

[0843] Step 10:

[0844] The server collects the execution results, adjusts them according to the user's emotions, and sends them to the terminal for feedback.

[0845] Step 11:

[0846] The device displays feedback sent from the server to the user, presenting it in a format that makes it easy for the user to understand the results.

[0847] This series of steps, through a process that takes user emotions into account, provides more appropriate and flexible data processing results.

[0848] (Example 2)

[0849] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0850] Current data processing systems can adequately analyze users' natural language input to generate optimal data processing workflows, but they struggle to respond to changes in user emotions and incorporate that feedback. As a result, users may have a less satisfying experience, and the system fails to meet expectations for interactive data processing.

[0851] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0852] In this invention, the server includes means for analyzing natural language input, means for recognizing user emotions, means for adjusting responses based on emotions, and means for selecting and generating workflows from a template library. This enables real-time responses to both user input and emotions, and provides optimized data processing workflows.

[0853] "Means of accepting input in natural language" refers to an interface that allows users to provide instructions and information to a system using natural language.

[0854] "Means of analyzing information and identifying basic information related to the subject" refers to the process of analyzing natural language input provided by the user and extracting basic information necessary for data processing from it.

[0855] "Means of querying users for missing information" refers to a mechanism that requests the user, in the form of questions, to complete the missing information identified as a result of the analysis.

[0856] "Means of understanding user emotions" refers to technologies that recognize emotional states from users' natural language input and responses.

[0857] "Means of adjusting response methods based on emotional information" refers to a process that modifies the system's responses and presentation methods according to the recognized emotions of the user, in order to provide an appropriate experience for the user.

[0858] "A means of selecting and combining appropriate workflows from a template library" refers to a method of constructing specific processing procedures by selecting the template that best suits the user's requirements from existing templates, or by combining multiple templates.

[0859] "A means of executing the generated workflow and providing feedback to the user" refers to the process of actually running the constructed data processing workflow and notifying the user of the processing results.

[0860] The system for carrying out this invention operates on a computer network including servers, terminals, and a user interface. Some processing utilizes a cloud-based platform.

[0861] The device provides a user interface, creating a space for users to input information using natural language. This interface allows input via text boxes or voice input, and utilizes tools such as the Google Cloud Natural Language API for natural language processing.

[0862] The server receives natural language input data sent from the terminal and analyzes the information via a natural language processing engine. This engine meticulously analyzes the input sentences and accurately extracts the basic information necessary for data processing. Simultaneously, the server uses sentiment recognition technology to understand the user's emotional state in real time from their input. Sentiment analysis utilizes technologies such as Microsoft Azure Text Analytics.

[0863] Based on emotional information, the server generates questions in an appropriate tone and content to ask the user for missing information, and presents them to the user through the terminal. By utilizing a generative AI model to adjust the prompt text as needed, the quality of communication is improved.

[0864] The user can respond to questions presented on the device and provide any necessary additional information. The emotion recognition engine also operates during this process, providing appropriate feedback based on the user's responses.

[0865] The server integrates all information obtained from the user and selects or generates the most appropriate workflow from the template library. This workflow is dynamically adjusted to match the user's needs based on the generated AI model.

[0866] The processing results obtained from the executed workflow are customized to the user's emotions and ultimately fed back to the user via the device. The results may also be displayed as infographics, providing a visually easy-to-understand format.

[0867] For example, if a user requests, "Find outliers in recent sales data," the system will analyze the input data and ask additional questions about the required data range and calculation criteria. Based on the user's response, it will generate and execute an appropriate workflow and provide feedback on the results in an appropriate format.

[0868] A concrete example of a prompt statement is, "Please tell me the data range required for this process." This allows for the provision of information that aligns with the user's needs and emotions.

[0869] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0870] Step 1:

[0871] The device provides a user interface, allowing users to input information and instructions in natural language. Users can use either a text box or voice input. The input data is captured by the device as natural language text.

[0872] Step 2:

[0873] The server receives natural language text from the terminal as input and sends it to the natural language processing engine. The natural language processing engine analyzes the text and extracts the basic information necessary for data processing. This process includes identifying key keywords and context.

[0874] Step 3:

[0875] Based on the analyzed information, the server uses an emotion recognition engine to recognize the user's emotional state. This input includes natural language text, and the output is either an emotion score or an emotion category. This allows for an understanding of the user's emotions.

[0876] Step 4:

[0877] The server identifies the missing information and generates a question in an appropriate tone, taking sentiment scores into consideration. Using a generative AI model, it creates a prompt and presents a specific question (e.g., "Please tell me the data range required for this process"). This question is then sent to the user.

[0878] Step 5:

[0879] The user enters their answers to the presented questions into the terminal. The additional information entered is specific response data to the questions. This allows for data completion.

[0880] Step 6:

[0881] The server retrieves additional information from the user and selects the most suitable workflow from the template library. Based on the retrieved additional information, data calculations and processing are performed, and the workflow is generated.

[0882] Step 7:

[0883] The server executes the generated workflow. This process performs each step for data processing, and in the process, user sentiment data is re-evaluated. The execution results are then obtained.

[0884] Step 8:

[0885] The server sends the final processing results to the terminal and notifies the user. The results are presented in a customized format based on the user's emotions. This output may take the form of an infographic or summary.

[0886] (Application Example 2)

[0887] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0888] Conventional natural language processing systems do not consider user emotions when analyzing input, and therefore do not optimize responses based on emotions. This can lead to decreased user satisfaction. Furthermore, fixed dialogue methods lack flexibility in sequential information supplementation, making it difficult to provide dynamic information that meets user needs.

[0889] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0890] In this invention, the server includes means for receiving input in natural language, means for analyzing information from the input natural language and identifying basic information related to the subject, and means for recognizing the user's emotions and adjusting the response content according to those emotions. This enables responses that are adapted to the user's emotions, improving the quality of interaction.

[0891] A "means of accepting natural language input" refers to an interface for receiving data and instructions from users in the form of natural language.

[0892] "Means for analyzing information and identifying basic information related to a subject" refers to the process of analyzing the content of natural language received from a user and extracting key points and basic information related to a specific subject based on that content.

[0893] "Means of querying users for missing information" refers to methods of identifying shortcomings in user instructions based on analysis results and asking users about them.

[0894] "Means of recognizing user emotions and adjusting responses accordingly" refers to a mechanism that analyzes the user's emotional state and dynamically changes the content of responses and information presented based on the results.

[0895] "A means of selecting and combining appropriate work procedures from a template library" refers to a system for selecting the most suitable predefined work procedures according to the situation and combining them as needed to construct new procedures.

[0896] "A means of executing the generated work procedure and providing feedback to the user on the results" refers to a function that actually carries out the created work procedure and reports the results and progress of that work to the user.

[0897] This invention relates to a system applied to a consumer robot that provides household support by utilizing natural language input from the user. The system components include a natural language processing engine (NLP engine), an emotion recognition engine, and a task scheduler. The hardware is a home robot equipped with a microphone for receiving voice input, a speaker for outputting responses, and a display for providing visual information.

[0898] The server first receives natural language input from the user through the robot's microphone. The received audio data is analyzed using an NLP engine to understand the user's request. At the same time, the emotion recognition engine is activated to recognize the user's emotional state in real time. This information is used to adjust the response in subsequent processes.

[0899] Once the analysis is complete, the server selects the optimal household chore procedure from the template library and, if necessary, combines and adjusts the order of tasks in a way that adapts to the user's emotions. The created procedure is then executed by a robot, and the progress and results are fed back to the user. The feedback re-evaluates the user's emotions and is reported in an appropriate format that reflects those emotions.

[0900] As a concrete example, suppose a user instructs a robot to "clean the kitchen next Friday." The NLP engine understands this instruction, and the emotion recognition engine detects that the user is tired. In this case, the robot proposes a more efficient cleaning plan and schedules the task in a way that reduces the user's burden.

[0901] Examples of prompts include, "How would the robot respond if the user said, 'When I've finished cleaning the kitchen, suggest a new recipe?'" and "How would you improve the suggestion if the user seemed dissatisfied?" In this way, the system enables efficient and adaptive responses based on the user's instructions and emotions.

[0902] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0903] Step 1:

[0904] The user inputs natural language instructions to the robot via a microphone. This input is voice data, which the terminal collects. This prepares the voice data for transmission to the server.

[0905] Step 2:

[0906] The server provides the audio data received from the terminal to the natural language processing engine, where it is converted into text-based instructions. This text data includes information about specific household tasks and desired actions. This conversion enables analysis in the next step.

[0907] Step 3:

[0908] The server uses the analysis capabilities of its natural language processing engine to extract important information from the text data. It identifies keywords and tasks related to the subject, and then uses an emotion recognition engine to evaluate the user's emotional state. At this stage, input data (text) and intermediate data (emotional information) are generated.

[0909] Step 4:

[0910] The server generates questions to prompt the user for additional information if any is missing. It adjusts the questions, including the appropriate tone and phrasing, taking sentiment recognition results into account. These questions are then sent to the terminal and displayed to the user, setting the stage for obtaining the user's response in the next step.

[0911] Step 5:

[0912] The user provides additional information in response to questions displayed on the device. This response is again collected as audio data and sent to the server. The user's emotional state is also monitored simultaneously.

[0913] Step 6:

[0914] The server integrates additional information and sentiment data received from the user, selects the optimal work procedure from the template library, and combines it as needed to generate a new workflow. This generated workflow is then ready to be executed by the task scheduler.

[0915] Step 7:

[0916] A workflow built by the server is executed by a robot. During the process, the user's emotions are continuously evaluated, and the workflow is dynamically adjusted based on that feedback. Finally, the results are reported to the user via a terminal. This report is presented in a format that reflects the user's emotions.

[0917] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0918] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0919] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0920] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0921] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0922] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0923] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0924] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0925] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0926] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0927] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0928] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0929] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0930] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0931] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0932] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0933] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0934] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0935] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0936] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0937] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0938] The following is further disclosed regarding the embodiments described above.

[0939] (Claim 1)

[0940] A means of accepting input in natural language,

[0941] A means for analyzing information from input natural language and identifying basic information related to the subject,

[0942] A means of querying the user for missing information based on the analysis results,

[0943] A means of selecting and combining appropriate workflows from a template library using additional information received from the user,

[0944] A means of executing the generated workflow and providing feedback to the user on the results,

[0945] A system that includes this.

[0946] (Claim 2)

[0947] The system according to claim 1, which utilizes generative artificial intelligence for natural language input analysis.

[0948] (Claim 3)

[0949] The system according to claim 1, comprising an interactive interface that sequentially supplements information using interaction with the user.

[0950] "Example 1"

[0951] (Claim 1)

[0952] A device that accepts natural language input,

[0953] A device that analyzes input natural language and identifies information processing procedures,

[0954] A device that requests additional information from the user based on the analysis results,

[0955] A device that selects and combines processing procedures to generate output based on additional information obtained from the user,

[0956] A device that executes the generated processing procedure and provides the execution result to the user via a display device,

[0957] An information processing system that includes this.

[0958] (Claim 2)

[0959] The information processing system according to claim 1, which uses a generative artificial intelligence model for natural language input analysis.

[0960] (Claim 3)

[0961] The information processing system according to claim 1, comprising an interactive user interface for sequentially supplementing information through dialogue.

[0962] "Application Example 1"

[0963] (Claim 1)

[0964] A means of accepting input in natural language,

[0965] A means for analyzing information from input natural language and identifying basic information related to the subject,

[0966] A means of inquiring with users about missing information based on the analysis results,

[0967] A means of selecting and combining appropriate work procedures from a set of templates using additional information received from the user,

[0968] A means of executing the generated work procedure and providing real-time feedback of the results to the user,

[0969] A means of analyzing data acquired in real time and utilizing it as optimized information for visualization in civic activities and business management,

[0970] A system that includes this.

[0971] (Claim 2)

[0972] The system according to claim 1, which uses generative artificial intelligence for natural language input analysis and performs dynamic updating and visualization of urban data.

[0973] (Claim 3)

[0974] The system according to claim 1, which features an interactive interface that sequentially supplements information using dialogue with users, and supports the optimization of urban operations.

[0975] "Example 2 of combining an emotion engine"

[0976] (Claim 1)

[0977] A means of accepting input in natural language,

[0978] A means for analyzing information from input natural language and identifying basic information related to the subject,

[0979] A means of querying the user for missing information based on the analysis results,

[0980] A means of understanding user emotions by analyzing user input using emotion recognition technology,

[0981] A means to adjust the system's response method based on emotional information and optimize the tone when querying for missing information,

[0982] A means of selecting and combining appropriate workflows from a template library using additional information received from the user,

[0983] A means of executing the generated workflow and providing feedback by customizing the results according to the user's emotional state,

[0984] A system that includes this.

[0985] (Claim 2)

[0986] The system according to claim 1, which utilizes an artificial intelligence model for natural language input analysis.

[0987] (Claim 3)

[0988] The system according to claim 1, comprising an interactive interface that sequentially supplements information using dialogue with the user and further adjusts the tone of the dialogue according to the user's emotions.

[0989] "Application example 2 when combining with an emotional engine"

[0990] (Claim 1)

[0991] A means of accepting input in natural language,

[0992] A means for analyzing information from input natural language and identifying basic information related to the subject,

[0993] A means of querying the user for missing information based on the analysis results,

[0994] A means of recognizing the user's emotions and adjusting the response accordingly,

[0995] A means of selecting and combining appropriate work procedures from a template library using additional information received from the user,

[0996] A means of executing the generated work procedure and providing feedback to the user on the results,

[0997] A system that includes this.

[0998] (Claim 2)

[0999] The system according to claim 1, which uses generative artificial intelligence for natural language input analysis and recognizes the user's emotions.

[1000] (Claim 3)

[1001] The system according to claim 1, comprising an interactive interface that sequentially supplements information using dialogue with the user and dynamically adjusts the content of the dialogue according to the user's emotions. [Explanation of Symbols]

[1002] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of accepting input in natural language, A means for analyzing information from input natural language and identifying basic information related to the subject, A means of inquiring with users about missing information based on the analysis results, A means of selecting and combining appropriate work procedures from a set of templates using additional information received from the user, A means of executing the generated work procedure and providing real-time feedback of the results to the user, A means of analyzing data acquired in real time and utilizing it as optimized information for visualization in civic activities and business management, A system that includes this.

2. The system according to claim 1, which uses generative artificial intelligence for natural language input analysis and performs dynamic updating and visualization of urban data.

3. The system according to claim 1, which includes an interactive interface that sequentially supplements information using dialogue with users, and supports the optimization of urban operations.