System
The system addresses language barriers by translating, grammar-checking, and summarizing multilingual information for rapid decision-making, enhancing communication efficiency in multilingual work environments.
Patent Information
- Application Number
- JP2024122742
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
Language barriers hinder communication efficiency and make it difficult to resolve business issues in multilingual work environments, particularly in centrally managing and reporting information to management.
A system that receives information in different languages, translates it into a common language, checks and corrects grammar, centrally manages the information, summarizes it using generative AI, and reports it to managers, enabling smooth communication and rapid decision-making.
Enables efficient information gathering, translation, grammar checking, data aggregation, and summary reporting across language barriers, supporting quick decision-making and smooth communication among employees.
Smart Images

Figure 2026021060000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In today's increasingly globalized world, when employees with different languages and cultures work together, language barriers hinder communication efficiency and make it difficult to resolve business issues. In particular, the process of centrally managing information and issues collected from multilingual employees, summarizing important content, and reporting it to management is cumbersome and inefficient. There is a need for a system that can solve this problem and support information sharing among employees and rapid decision-making. [Means for solving the problem]
[0005] This invention provides a means to receive information entered in different languages, translate it into a common language, and check and correct the grammar of the translated information.It also provides a means to centrally manage multiple pieces of information, summarize that information using generation AI, and ultimately report it to managers, thereby eliminating communication barriers between employees and enabling the rapid consolidation and summarization of work-related issues and smooth reporting to management.
[0006] 1. "Different languages" refers to languages used in multiple different countries or regions.
[0007] 2. A "lingua franca" refers to a standard language used by people who speak different languages to facilitate the sharing of information.
[0008] 3. "Translation" refers to the process of converting text or words written in one language into another.
[0009] 4. "Grammar checking" refers to the process of verifying whether a sentence is grammatically correct and correcting any errors.
[0010] 5. "Centralized information management" refers to consolidating and managing multiple pieces of information in one place.
[0011] 6. "Generative AI" refers to systems that use artificial intelligence techniques to generate new information or summaries.
[0012] 7. "Summarizing" refers to extracting important points from multiple pieces of information and summarizing them concisely.
[0013] 8. "Manager" refers to a person or position with the authority to make important decisions within a company or organization.
[0014] 9. "Means for receiving" refers to the functions and devices for receiving information.
[0015] 10. "Reporting means" means any facility or device that communicates aggregated and summarized information to another person or system. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] As a concrete example of the present invention, we will explain a multilingual collaborative task aggregation AI service. This system consists of three main components: a server, a terminal, and a user, each of which has a specific function.
[0038] Handling User Input
[0039] Users input their business issues into the terminal in their native language. This input is done through a text field, and after the user has written the issue, they press the "Submit" button. This operation causes the terminal to send the user's input as data to the server.
[0040] Translation and Grammar Check
[0041] The server receives user input sent from the terminal. The received data is first passed to a multilingual translation AI module, which translates the input language into a common language (e.g., English). The translated text is then sent to a grammar checker module, where it is checked for grammatical errors and corrected if necessary.
[0042] Aggregation and Summarization
[0043] The server centralizes translated assignments collected from multiple users. This process involves aggregating each user's input into a database. The aggregated data is then passed to a generative AI module, which further extracts and summarizes key points. This generative AI module has learned from past data and patterns to efficiently and accurately summarize information.
[0044] Reporting to management
[0045] The server sends the generated summary to an administrator interface, which updates in real time, allowing the administrator to instantly view the summarized information and make quick decisions based on the summary.
[0046] Specific examples
[0047] If a user inputs the issue "Customer feedback is delayed" in Japanese, the input is sent to the server via the terminal. The server translates this input into English as "Customer feedback is delayed." The grammar checker then corrects it to make it grammatically correct. Next, multiple similar reports are aggregated and a generative AI creates a summary: "Several employees report delays in customer feedback." Finally, this summary is sent to the administrator interface, where it is reviewed and evaluated by the administrator.
[0048] In this way, the present invention helps multilingual employees efficiently gather information, translate, check grammar, aggregate data, and summarize it, and quickly report it to managers, thereby supporting smooth communication and quick decision-making across language barriers.
[0049] The processing flow will be explained below.
[0050] Step 1:
[0051] The user enters business issues and information into the terminal in their native language. After selecting the language they want to use and entering the issue in the text field, they press the "Submit" button.
[0052] Step 2:
[0053] The terminal sends the user's input to the server. The terminal sends the input data to the server as an HTTP request. This data includes the language information selected by the user and the assignment text.
[0054] Step 3:
[0055] The server passes the received information to a multilingual translation AI, which translates it into a common language (e.g., English). The server analyzes the received data, sends the text and language information to the translation API, and obtains the translation results.
[0056] Step 4:
[0057] The server passes the translated text to the grammar checker module to check and correct grammatical errors. The server sends the translated text to the grammar checker API to make it grammatically correct.
[0058] Step 5:
[0059] The server centralizes the translated assignments from all employees. The server stores the translated and grammar-checked text in a database and aggregates input from different employees.
[0060] Step 6:
[0061] The server passes the aggregated data to the generative AI to create a summary. The server extracts multiple issues stored in the database and passes them to the generative AI module to create a summary.
[0062] Step 7:
[0063] The server sends the generated summary to the administrator interface, where it displays the summary created by the generating AI in real time, allowing for immediate confirmation.
[0064] Step 8:
[0065] The administrator checks the summary on the management screen and makes a decision. The administrator evaluates the summary content through the interface and takes necessary action or makes a decision.
[0066] Through specific actions at each step, the entire system functions efficiently, supporting information sharing and quick decision-making among employees who speak different languages.
[0067] Example 1
[0068] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0069] In today's global companies, when employees who speak different languages work together, language barriers become an obstacle, making it difficult to communicate efficiently and make quick decisions. To address this issue, systems are needed that can efficiently translate information, check grammar, aggregate data, and summarize it. In existing systems, these processes are often performed manually, requiring time and effort. Furthermore, a lack of centralized information management makes it difficult for managers to quickly obtain the information they need.
[0070] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0071] In this invention, the server includes means for receiving information entered in different languages, means for translating the received information into a common language, means for checking and correcting the grammar of the translated information, means for centrally managing multiple pieces of information, means for summarizing the centrally managed information using a generative model, and means for reporting the summarized information to an administrator in real time. This enables smooth communication across language barriers and efficient aggregation and summarization of information. As a result, administrators can quickly obtain the information they need and make quick decisions.
[0072] "Means for receiving" refers to the function for importing information entered in different languages from a terminal into a server.
[0073] "Means for translating into a common language" refers to a function for converting received information from a different language into one pre-specified language (primarily English).
[0074] "Grammar checking and correction means" refers to functionality for detecting and, if necessary, correcting grammatical errors in the translated information.
[0075] "Means for centralized management" refers to a function for aggregating information received from multiple users into a central database and managing it in a unified manner.
[0076] "Means for summarizing using a generative model" refers to a function that uses a generative AI model based on aggregated information to extract important points and create a summary.
[0077] "Means of reporting in real time" refers to a function that allows the administrator to be notified of the generated summary immediately without any time lag.
[0078] MODE FOR CARRYING OUT THE INVENTION
[0079] This invention is a system for a multilingual collaborative task aggregation AI service. This system consists of three elements: a server, a terminal, and a user. Each element has a specific function and works together to solve business tasks.
[0080] Handling User Input
[0081] Users input business issues into the terminal in their native language. This input is done through a text field, and after the user describes the issue, they press the "Submit" button. This operation causes the terminal to generate data and send it to the server as JSON-formatted data. For example, if a user inputs "Customer feedback is delayed," this text is sent from the terminal to the server.
[0082] Translation Processing
[0083] The server receives the data sent from the device. It analyzes the received data and passes it to a multilingual translation AI module. This translation AI module may use, for example, the Google Translate API or the DeepL API. The input Japanese text is translated into English, the common language. For example, "Customer feedback is delayed" is translated to "Customer feedback is delayed."
[0084] Grammar Check
[0085] The translated text is sent to a grammar checker module, which can use the Grammarly API or LanguageTool API, for example. The grammar checker detects grammatical errors and corrects them if necessary. For example, "Customer feedback is delayed." is checked for grammar and any necessary corrections are made.
[0086] Data Aggregation
[0087] The server aggregates all the translated texts sent by users into a database. This process uses MySQL or PostgreSQL as a database management system (DBMS), which allows for centralized management of each user's input.
[0088] Data Summary
[0089] The aggregated data is passed to a generative AI module, which uses generative AI models such as OpenAI's GPT-3 and ChatGPT. The generative AI learns from past data and patterns, extracts key points, and summarizes the text. For example, if multiple users submit similar reports, the generative AI will generate a summary such as "Several employees report delays in customer feedback."
[0090] Reporting to management
[0091] The generated summary is sent to an administrator interface, where a dashboard is displayed that updates in real time, allowing administrators to instantly view the summarized information. For example, when an administrator opens the dashboard, they might see a summary that reads, "Several employees report delays in customer feedback."
[0092] Example prompt sentences
[0093] Below are some specific examples of prompt sentences to input to the generative AI model.
[0094] 1. The user types "Customer feedback is delayed" in Japanese into the device and presses the send button.
[0095] 2. The device sends this input to the server as JSON format data.
[0096] 3. The server passes the received data to the Google Translate API and translates it into English: "Customer feedback is delayed."
[0097] 4. Check the grammar of the translated text using the Grammarly API and make corrections if necessary.
[0098] 5. The server stores and aggregates all user data in a MySQL database.
[0099] 6. Pass the aggregated data to ChatGPT and generate a summary: “Several employees report delays in customer feedback.”
[0100] 7. The server sends a summary to the administrator's dashboard, where the administrator can instantly view it.
[0101] In this way, the present invention enables smooth communication in a multilingual environment, and supports efficient information gathering and summarization, as well as rapid decision-making.
[0102] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0103] Step 1:
[0104] The user inputs the issue into the terminal. The user writes the business issue in their native language into the text field on the terminal and presses the "Submit" button. This operation converts the entered text into JSON format by the terminal and sends it to the server. The input is the text "Customer feedback is delayed," and the output is JSON format data.
[0105] Step 2:
[0106] The server receives the data sent from the device. It analyzes the received JSON data and passes it to a multilingual translation AI module. This translation AI module uses, for example, a multilingual translation API. It receives JSON data as input and obtains translated text data as output. For example, the Japanese data "Customer feedback is delayed" is translated into the English data "Customer feedback is delayed."
[0107] Step 3:
[0108] The server sends the translated text to a grammar checker module. The grammar checker module uses, for example, a grammar check API. It receives the translated text data as input, detects grammatical errors, and outputs corrected text data as necessary. For example, the text "Customer feedback is delayed." is grammatically checked to ensure there are no errors.
[0109] Step 4:
[0110] The server aggregates all translated text submitted by users into a database. The database management system (DBMS) can be, for example, a relational database. It receives translated text data as input and outputs it as a centralized database record. For example, text submitted by multiple users can be aggregated into a single database.
[0111] Step 5:
[0112] The server passes the aggregated data to a generative AI module. This module uses a generative AI model that takes the aggregated data from the database as input, extracts key points, and outputs summarized text data. For example, multiple matching reports might be summarized as "Several employees report delays in customer feedback."
[0113] Step 6:
[0114] The server sends the generated summary to an administrator interface, which displays a dashboard that updates in real time, allowing administrators to instantly view the summarized information. The interface receives summarized text data as input and displays it as an administrator-accessible dashboard as output. For example, a real-time summary such as "Several employees report delays in customer feedback" appears on the administrator's dashboard.
[0115] (Application example 1)
[0116] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0117] In an environment where employees speak a wide variety of languages, it is necessary to quickly and efficiently collect, translate, and summarize business issues and provide accurate information to managers. Real-time problem reporting and rapid decision-making are particularly important in logistics centers. However, the overlap of different languages and manual information processing can lead to delays and errors in the collection, translation, and transmission of information.
[0118] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0119] In this invention, the server includes a means for receiving information input in different languages, a means for translating the received information into a common language, and a means for checking and correcting the grammar of the translated information. This provides a means for centrally managing multiple pieces of information, a means for summarizing the centrally managed information, and a means for reporting the summarized information to an administrator. This allows for the input of information in real time using a smart device, and for the input information to be translated, checked for grammar, and summarized, and then reported promptly to an administrator.
[0120] "Information entered in different languages" refers to text data entered by users in their various native languages.
[0121] A "common language" is a unified language (e.g., English) that is translated from different languages.
[0122] "Grammar checking and correction" is the process of making the translated text grammatically correct.
[0123] "Centralized management" refers to the process of aggregating multiple pieces of information into a central database and managing them in an integrated manner.
[0124] "Summarizing" refers to extracting important points from aggregated information and summarizing them concisely.
[0125] "Reporting to administrator" is the process of sending the processed information to an administrator interface so that the administrator can review it.
[0126] "Inputting information in real time using a smart device" refers to inputting information instantly from a mobile device such as a smartphone or tablet.
[0127] A "generative model" is an algorithm that uses machine learning techniques to summarize and generate information.
[0128] This paper describes a system that realizes a multilingual collaborative task aggregation AI service in a logistics center. This system consists of three main elements: a server, a terminal, and a user, each of which has a specific function.
[0129] Handling User Input
[0130] Users input their work-related tasks in their native language using a smartphone. This input is done through a text field, and after the user has written the task, they press the "Submit" button to complete the task. The device then sends the user's input as data to the server.
[0131] Translation and Grammar Check
[0132] The server receives user input sent from the device. The received data is first translated into a common language (e.g., English) using the Google Translator API. This translated text is then sent to the grammar checking module via OpenAI's API, where grammatical errors are checked and corrected.
[0133] Aggregation and Summarization
[0134] The server centrally manages translated assignments collected from multiple users and aggregates them into a database. The aggregated data is then extracted and summarized by OpenAI's generative AI model. This generative AI model learns from past data and patterns to efficiently and accurately summarize information.
[0135] Report to the administrator
[0136] The server sends the generated summary to an administrator interface, which updates in real time, allowing administrators to instantly see the summarized information and make quick decisions.
[0137] Specific examples
[0138] If a user inputs an issue in Japanese such as "Errors in the inventory management system occur frequently," the input is sent to the server via the terminal. The server translates this input into English using the Google Translator API, resulting in "Errors in the inventory management system occur frequently." The translated text is then checked for grammar using the OpenAI API, and corrected to "Errors in the inventory management system are occurring frequently." The generation AI generates a summary from many similar reports: "Frequent errors in inventory system." This summary is sent to the administrator interface, where the administrator can take prompt action based on it.
[0139] Example prompt sentence:
[0140] Grammar check prompt: "Please correct the following text: Errors in the inventory management system occur frequently."
[0141] Summary prompt: "Summarize the following text: Errors in the inventory management system are occurring frequently."
[0142] In this way, the present invention enables multilingual employees to efficiently gather information, translate, check grammar, aggregate data, and summarize, and quickly report to management, thereby supporting smooth communication and quick decision-making across language barriers in logistics centers.
[0143] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0144] Step 1:
[0145] A user uses a smart device to input a business issue in their native language. Specifically, the user writes the issue in a text field on the smartphone application and presses the "Submit" button. The input issue (e.g., "Errors are occurring frequently in the inventory management system") is sent from the device to the server. The input data is text data in the user's native language.
[0146] Step 2:
[0147] The server receives user input sent from the device. The received data is first sent to the Google Translator API, where the native text data is translated into a common language (e.g., English). The translated text data (e.g., "Errors in the inventory management system occur frequently.") is generated.
[0148] Step 3:
[0149] The server sends the translated text data to OpenAI's API, where a grammar check is performed using a generative AI model. Specifically, the server detects and corrects grammatical errors in the translated text, generating corrected text data (e.g., "Errors in the inventory management system are occurring frequently.").
[0150] Step 4:
[0151] The server stores the grammar-checked text data in a database and manages it centrally. Multiple similar data sets are aggregated, and the text data is stored in a management database. This ensures that assignments from multiple users are managed in an integrated manner.
[0152] Step 5:
[0153] The server sends the aggregated data to OpenAI's generative AI model, which extracts key points and creates a summary. Specifically, the server selects key content from multiple translated and corrected text data and generates a summary text (e.g., "Frequent errors in inventory system.").
[0154] Step 6:
[0155] The server sends the generated summary text to an administrator interface, where administrators can view the summarized information using an interface that updates in real time, providing administrators with accurate information for quick decision making.
[0156] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0157] The present invention is a system that incorporates an emotion engine that recognizes the user's emotions, as well as a means for receiving information entered in different languages, translating it into a common language, and checking and correcting the grammar of the information. This system consists of three main components: a server, a terminal, and a user, each of which has a specific function.
[0158] User Input Processing and Emotion Recognition
[0159] Users input work-related issues and information into the device in their native language. After selecting the language they want to use and entering the issue in the text field, they press the "Submit" button. The device then sends this input data to the server. During this process, the device also collects emotional information along with the user's input data. The emotional information is recognized by an emotion engine that analyzes the user's input content and behavior during input.
[0160] Translation and Grammar Check
[0161] The server receives user input and emotion information sent from the device. The received data is first passed to a multilingual translation AI module, which translates the input language into a common language (e.g., English). The translated text is then sent to a grammar checker module, where it is checked for grammatical errors and corrected if necessary.
[0162] Unified management of emotional information
[0163] The server centrally manages the emotional information along with the translated and grammar-checked information. The server stores this information in a database and aggregates input data from different employees. At this time, the emotional information is treated like regular text data and stored in the database.
[0164] Aggregation and Summarization
[0165] The server aggregates multiple issues stored in the database and passes them to the generation AI to create a summary. Emotional information is also provided to the generation AI, and the tendency and intensity of emotions are reflected in the summarization process. For example, if many employees are feeling stressed, that emotional information will be noted as part of the summary.
[0166] Report to the administrator
[0167] The server sends the generated summary and emotion information to the administrator interface, which is updated in real time, allowing the administrator to instantly check the summary and emotion information. This allows the administrator to understand not only the text information but also the emotional state of employees, enabling them to make better decisions.
[0168] Specific examples
[0169] If a user inputs an issue in Japanese, such as "Customer feedback is delayed. I am very upset," and the input includes emotional expressions, the input is sent to the server via the terminal. The server translates this into English as "Customer feedback is delayed. I am very upset." A grammar checker then corrects the input to make it grammatically correct. The server also uses an emotion engine to recognize emotions such as "confused" and "stressed." Next, multiple similar reports are aggregated, and a generative AI creates a summary: "Several employees report delays in customer feedback and are very stressed." Finally, this summary and emotional information are sent to the administrator interface, where the administrator reviews it and considers countermeasures.
[0170] This system integrates multilingual information collection, translation, grammar checking, and emotion recognition, and can centrally manage and summarize data for quick reporting, eliminating communication barriers caused by language and cultural differences and supporting appropriate decision-making that takes into account the emotional state of employees.
[0171] The processing flow will be explained below.
[0172] Step 1:
[0173] The user enters work-related issues and information into the device in their native language. After selecting the language they want to use and entering the issue in the text field, they press the "Submit" button. This operation also collects emotional information.
[0174] Step 2:
[0175] The device sends the user's input and emotional information to the server. The input content and emotional information are sent to the server as an HTTP request. The emotional information is the result of an emotion engine analyzing the user's input content and behavior at the time of input.
[0176] Step 3:
[0177] The server passes the received information to a multilingual translation AI, which translates it into a common language (e.g., English). The server analyzes the received data, sends the text and language information to the translation API, and obtains the translation results.
[0178] Step 4:
[0179] The server passes the translated text to the grammar checker module, which checks and corrects grammatical errors, and the translation result is sent to the grammar checker API, which formats it to be grammatically correct.
[0180] Step 5:
[0181] The server centralizes the translated and grammar-checked information, as well as the sentiment information, which is then stored in a database that aggregates input from different employees.
[0182] Step 6:
[0183] The server passes the collected data to the generation AI to create a summary. Using the issue information and emotion information stored in the database, the generation AI module extracts important points and reflects them in the summary.
[0184] Step 7:
[0185] The server sends the generated summary and emotion information to the administrator interface, which is updated in real time, allowing the administrator to instantly check the summary and emotion information.
[0186] Step 8:
[0187] The administrator checks the summary on the management screen and makes a decision. The administrator evaluates the summary content and sentiment information through the interface and quickly takes necessary actions and makes decisions.
[0188] As a specific example, if a user inputs and submits the issue in Japanese, "Customer feedback is delayed. I am very upset.", the input and emotions are recognized as "confused" and "stressed." The server translates the received input into English, resulting in "Customer feedback is delayed. I am very upset." The grammar checker corrects the input to make it grammatically correct, and then stores it in the database. Next, it is aggregated with similar input from other employees, and the generative AI creates a summary: "Several employees report delays in customer feedback and are very stressed." This summary and emotional information are sent to the administrator interface, where administrators can make quick decisions based on it.
[0189] In this way, the entire system functions efficiently, supporting information sharing among employees who speak different languages and quick decision-making that takes emotional information into account.
[0190] Example 2
[0191] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0192] There is a need for a system that can efficiently collect, translate, check grammar, and analyze emotions from users who speak different languages, and then report the results to administrators. This system must also be able to accurately analyze users' emotional information and generate summaries that take this information into account, thereby solving the problem of providing administrators with information that includes the users' emotional state.
[0193] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0194] In this invention, the server includes means for receiving information input in different languages, means for translating the received information into a common language, means for checking and correcting the grammar of the translated information, means for centrally managing multiple pieces of information, means for summarizing the centrally managed information, means for reporting the summarized information to an administrator, means for collecting emotional information from the input information, means for analyzing the emotional information, and means for generating a summary that includes the emotional information. This makes it possible to efficiently collect and manage information from users who speak different languages and to provide information that includes their emotional states.
[0195] "Different languages" refers to languages used by users in multiple different countries or regions.
[0196] "Input information" refers to text data and assignments entered by the user using the terminal.
[0197] "Means of receiving" refers to the mechanism or method by which the server receives data sent from the terminal.
[0198] "Common language" refers to a standard language that is used uniformly within the system, such as English.
[0199] "Means of translation" refers to the functions and methods for converting information entered in different languages into a common language.
[0200] "Grammar checking and correction means" refers to a method or module for detecting grammatical errors in the translated text and correcting it to make it grammatically correct.
[0201] "Means of centralized management" refers to a method of integrating multiple input data and translation results and centrally managing them in a database or management system.
[0202] "Means of summarization" refers to techniques or methods, such as the use of generative AI models, to succinctly summarize large amounts of input information or data.
[0203] "Means for reporting to administrator" refers to the method or interface for notifying the administrator of the summarized information.
[0204] "Emotional information" refers to emotional states analyzed from information entered by the user, such as data on "confusion" or "stress."
[0205] "Means for collecting emotional information" refers to the technology or method for analyzing user input data to recognize and collect emotions.
[0206] "Means for analyzing emotional information" refers to technology for analyzing collected emotional data and evaluating its content and trends.
[0207] "Means for generating summaries that include emotional information" refers to techniques and methods for creating summaries of input data that take emotional information into account.
[0208] The present invention is a system that receives information entered in different languages, translates it into a common language, and checks and corrects the grammar of the information. This system also has the function of collecting and analyzing emotional information, centrally managing the collected information, and finally summarizing and reporting it to an administrator. Specific embodiments of this system are described below.
[0209] Users input business issues and information into the device in their native language. For example, they can enter something like "Feedback from customers is delayed. We are having a lot of trouble" in Japanese and press the "Send" button. Users can first select the language they want to use and then enter the issue or information in the text field.
[0210] The device receives the user's input data. At the same time, a built-in emotion engine simultaneously collects emotional information based on the input content and the user's input speed. This emotion engine analyzes and recognizes emotions such as "confusion" or "stress" from the user's text content.
[0211] The server receives the user's input data and emotional information sent from the device. The received data is first passed to a multilingual translation AI module, which translates the different input languages (such as Japanese) into a common language (such as English). An industry-standard AI translation engine is often used as the translation module.
[0212] The translated text is then sent to a grammar checker module, which can incorporate existing grammar checking software such as Grammarly, to check for grammatical errors and correct them if necessary.
[0213] Once the translation and grammar check are complete, the information is centrally managed by a server. This information is stored in a database that aggregates input data from different employees. This database management system may use MySQL or PostgreSQL.
[0214] Next, the server aggregates multiple issues stored in the database and passes them to a generation AI to create a summary. Here, emotional information is also provided to the generation AI, and the tendency and intensity of emotions are reflected in the summarization process. For example, a summary is generated that includes emotional information, such as "Most employees are stressed by delays in feedback from customers."
[0215] Finally, the server sends the generated summary and emotion information to the administrator interface, which is updated in real time, allowing the administrator to instantly check the information. This allows the administrator to understand the employee's emotional state and provides information for better decision-making.
[0216] For example, if a user inputs an issue in Japanese such as "Customer feedback is delayed. I am very upset." and the input contains emotional expressions, the input is sent to the server via the terminal. The server translates this into English as "Customer feedback is delayed. I am very upset." A grammar checker then corrects the input to make it grammatically correct. The server also uses an emotion engine to recognize emotions such as "confused" and "stressed." Next, multiple similar reports are aggregated, and a generative AI creates a summary such as "Several employees report delays in customer feedback and are very stressed." Finally, this summary and emotional information are sent to the administrator interface, where the administrator reviews it and considers countermeasures.
[0217] An example prompt is, "Recognize the sentiment from the user's input data and summarize the text in a way that includes sentiment. Summarize the text as follows: 'Several employees report delays in customer feedback and are very stressed.'"
[0218] As described above, the present invention can provide a system that can efficiently manage information regardless of language differences and can report information including emotional information to a manager.
[0219] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0220] Step 1:
[0221] The user inputs the issue or information into the device in their native language. After inputting, they select the language they want to use, enter something like "Feedback from customers is delayed. We are having a lot of trouble," in the text field, and press the "Send" button. This saves the input information as text data on the device.
[0222] Step 2:
[0223] The device receives text data entered by the user. At the same time, the emotion engine analyzes the text content and behavioral data such as the user's typing speed, and collects and recognizes emotional information such as "confusion" or "stress." The input is text data, and the output is text data and emotional information. Specifically, the device also collects the user's keystroke data and sends it to the emotion engine for analysis.
[0224] Step 3:
[0225] The device sends the user's text data and emotional information to the server. Specifically, the device requests the server to send data using the HTTP protocol. The input is text data and emotional information, and the output is sent to the server.
[0226] Step 4:
[0227] The server receives text data and emotion information sent from the device. The received data is first passed to a multilingual translation AI module, which translates the input language (e.g., Japanese) into a common language (e.g., English). The input is Japanese text data, and the output is English text data. The server calls the translation engine to convert the data.
[0228] Step 5:
[0229] The server sends the translated text data to the grammar checker module, which checks and corrects grammatical errors. The input is English text data, and the output is English text data with corrected grammar. The server passes the data to the grammar checker module and receives the corrected results.
[0230] Step 6:
[0231] The server stores the corrected text data and emotion information in a database for centralized management. The input is the corrected text data and emotion information, and the output is saving to the database. The server inserts data into the database management system using SQL queries.
[0232] Step 7:
[0233] The server aggregates multiple issues and emotion information stored in the database and passes it to the generative AI model to create a summary. The input is multiple data in the database, and the output is summary data. Specifically, the server retrieves the data using an SQL query and provides it to the generative AI model along with a prompt statement to generate a summary. An example of a prompt statement is "Several employees report delays in customer feedback and are very stressed."
[0234] Step 8:
[0235] The server sends the generated summary and emotion information to the administrator interface. The input is summary data and emotion information, and the output is interface updates. The server sends data to the administrator's browser in real time using WebSocket or HTTP protocol and updates the interface.
[0236] Through these steps, the system can efficiently translate, check grammar, analyze sentiment, and aggregate information in different languages to generate summaries and provide them to administrators.
[0237] (Application example 2)
[0238] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0239] In the past, when communicating between different languages, language translation, grammar checking, and emotion recognition tended to be performed in separate systems, making it difficult to centrally manage and summarize information and report it to managers immediately.In addition, in physical stores, there was a lack of means to receive feedback from multilingual customers in real time and process it smoothly, which placed a heavy burden on staff.
[0240] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving information input in different languages, means for translating the received information into a common language, means for checking and correcting the grammar of the translated information, means for recognizing the user's emotions, means for centrally managing multiple pieces of information, means for summarizing the centrally managed information, means for reporting the summarized information to a manager, and components that allow store staff to input customer feedback on-site in real time and instantly check the translation results and emotional information. This centralizes feedback processing in a multilingual environment, making it possible to simultaneously manage customer feedback and the emotional state of staff.
[0241] The "means for receiving information input in a different language" is a function for receiving information input by a user in a different language at a terminal.
[0242] The "means for translating received information into a common language" is a function for automatically translating received information in multiple languages into a common language.
[0243] The "means for checking and correcting the grammar of translated information" is a function for detecting grammatical errors in translated information and automatically correcting them.
[0244] The "means for recognizing the user's emotions" is a function that analyzes the user's input content and behavior at the time of input and identifies the user's emotional state.
[0245] "Means for centrally managing multiple pieces of information" refers to a function that consolidates and manages multiple pieces of input information and related emotional information in one place.
[0246] "Means for summarizing centrally managed information" is a function that generates a short summary of the main points based on the aggregated information.
[0247] The "means for reporting summarized information to the administrator" is a function for instantly notifying the administrator of summarized information and emotional information.
[0248] "A component that allows store staff to input customer feedback in real time on-site and instantly check the translation results and emotional information" is a device that allows store staff to input feedback from customers in real time and instantly view it along with the translation and emotional analysis results.
[0249] This invention is a unified management system that receives customer feedback in real time in a multilingual environment in a brick-and-mortar store, translates it, checks grammar, and recognizes emotions. Specific embodiments of this system are described below.
[0250] System configuration
[0251] The system mainly consists of three main components: a server, a terminal, and a user.
[0252] Server: Translates data, recognizes emotions, checks grammar, and centralizes and summarizes information.
[0253] Terminal: A device where store clerks can input customer feedback in real time and check the translated results and sentiment analysis results. This can be a smartphone, smart glasses, or a head-mounted display.
[0254] Users: Includes customers and store clerks, where customers provide feedback and store clerks input it into the terminal.
[0255] Program Overview
[0256] The server receives information (feedback) entered in different languages and translates it into a common language. Translation services such as the Google Cloud Translation API are used for translation. The translated text is then checked for grammar using services such as the Grammarly API and corrected as necessary. Emotion recognition is performed using services such as the Microsoft Azure Emotion API to analyze the user's emotional state.
[0257] The received information and analyzed sentiment information are centrally managed on a server. Multiple feedback information is stored in a database and summarized using a generative AI model, allowing managers to efficiently grasp the key points.
[0258] Hardware and software used
[0259] Hardware:
[0260] Smartphone
[0261] Smart Glasses
[0262] head-mounted display
[0263] software:
[0264] Google Cloud Translation API (translation)
[0265] Microsoft Azure Emotion API (emotion recognition)
[0266] Grammarly API (Grammar Check)
[0267] Processing Flow
[0268] 1. User (store clerk): Enters customer feedback into a device such as a smartphone.
[0269] 2. Terminal: Sends the entered information to the server.
[0270] 3. Server:
[0271] Translate information.
[0272] Check and correct the grammar of the translated text.
[0273] Recognize user emotions.
[0274] 4. Server: Centralizes the translated and analyzed information, generates summaries, and reports to the administrator.
[0275] 5. Terminal: Displays translation results and sentiment analysis results in real time.
[0276] Specific examples
[0277] A customer provides feedback in Japanese, saying, "Customer feedback is delayed. I am very upset." The store clerk enters this into a terminal. This input is sent to the server, where it is translated into English, becoming, "Customer feedback is delayed. I am very upset." Grammar checks are then performed to ensure the correct form is achieved. Emotion recognition identifies emotions such as "confusion" and "stress," and the aggregated data generates a summary: "Several employees report delays in customer feedback and are very stressed." Finally, this summary and emotional information are reported to management.
[0278] Prompt Sentence Examples
[0279] "Please translate the feedback below, check for grammar, and recognize emotional information.
[0280] Language: Japanese
[0281] Feedback: "Customers are very unhappy"
[0282] As described above, it is possible to simultaneously manage customer feedback and the emotional state of staff in a multilingual environment, and a system is provided that can instantly translate and respond to communication between store staff and customers.
[0283] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0284] Step 1:
[0285] User (store clerk): The store clerk uses a device such as a smartphone or smart glasses to input customer feedback in real time. The input data is in text format, for example, "Customer feedback is delayed. We are very troubled." The input data is temporarily stored in the device.
[0286] Step 2:
[0287] Terminal: The terminal sends the input text data to the server, along with input language information (e.g., feedback in Japanese and its timestamp). A secure communication protocol (e.g., HTTPS) is used for this data transmission.
[0288] Step 3:
[0289] Server: The server receives the text data and first translates it into a common language (e.g., English) using the Google Cloud Translation API. The input is the original Japanese feedback, and the output is the translated data, such as "Customer feedback is delayed. I am very upset."
[0290] Step 4:
[0291] Server: After getting the translated data, we use the Grammarly API to perform a grammar check. Here, the input is the translated English text and the output is the grammatically corrected text. The corrections could be, for example, "Customer feedback has been delayed. I am very upset."
[0292] Step 5:
[0293] Server: After the grammar check is complete, emotion recognition is performed using the Microsoft Azure Emotion API. The input is the user's text data ("Customer feedback has been delayed. I am very upset."), and the output is the recognized emotion information (e.g., "confused" or "stressed").
[0294] Step 6:
[0295] Server: The data that has been translated, checked for grammar, and recognized for emotion is managed in a centralized database. The server stores this data and also assigns metadata such as timestamps to each piece of data. The input is the original feedback and various analyzed data, and the output is a centralized database entry.
[0296] Step 7:
[0297] Server: Multiple pieces of feedback stored in a database are summarized using a generative AI model. The input is a set of accumulated database entries, and the output is summary information, such as "Several employees report delays in customer feedback and are very stressed."
[0298] Step 8:
[0299] Server: Finally, the generated summary information and related sentiment information are sent to the administrator interface, allowing the administrator to check information based on staff status and customer feedback in real time. The input is the summarized data, and the output is the display results on the interface accessed by the administrator.
[0300] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0301] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0302] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0303] [Second embodiment]
[0304] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0305] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0306] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0307] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0308] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0309] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0310] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0311] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0312] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0313] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0314] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0315] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0316] As a concrete example of the present invention, we will explain a multilingual collaborative task aggregation AI service. This system consists of three main components: a server, a terminal, and a user, each of which has a specific function.
[0317] Handling User Input
[0318] Users input their business issues into the terminal in their native language. This input is done through a text field, and after the user has written the issue, they press the "Submit" button. This operation causes the terminal to send the user's input as data to the server.
[0319] Translation and Grammar Check
[0320] The server receives user input sent from the terminal. The received data is first passed to a multilingual translation AI module, which translates the input language into a common language (e.g., English). The translated text is then sent to a grammar checker module, where it is checked for grammatical errors and corrected if necessary.
[0321] Aggregation and Summarization
[0322] The server centralizes translated assignments collected from multiple users. This process involves aggregating each user's input into a database. The aggregated data is then passed to a generative AI module, which further extracts and summarizes key points. This generative AI module has learned from past data and patterns to efficiently and accurately summarize information.
[0323] Reporting to management
[0324] The server sends the generated summary to an administrator interface, which updates in real time, allowing the administrator to instantly view the summarized information and make quick decisions based on the summary.
[0325] Specific examples
[0326] If a user inputs the issue "Customer feedback is delayed" in Japanese, the input is sent to the server via the terminal. The server translates this input into English as "Customer feedback is delayed." The grammar checker then corrects it to make it grammatically correct. Next, multiple similar reports are aggregated and a generative AI creates a summary: "Several employees report delays in customer feedback." Finally, this summary is sent to the administrator interface, where it is reviewed and evaluated by the administrator.
[0327] In this way, the present invention helps multilingual employees efficiently gather information, translate, check grammar, aggregate data, and summarize it, and quickly report it to managers, thereby supporting smooth communication and quick decision-making across language barriers.
[0328] The processing flow will be explained below.
[0329] Step 1:
[0330] The user enters business issues and information into the terminal in their native language. After selecting the language they want to use and entering the issue in the text field, they press the "Submit" button.
[0331] Step 2:
[0332] The terminal sends the user's input to the server. The terminal sends the input data to the server as an HTTP request. This data includes the language information selected by the user and the assignment text.
[0333] Step 3:
[0334] The server passes the received information to a multilingual translation AI, which translates it into a common language (e.g., English). The server analyzes the received data, sends the text and language information to the translation API, and obtains the translation results.
[0335] Step 4:
[0336] The server passes the translated text to the grammar checker module to check and correct grammatical errors. The server sends the translated text to the grammar checker API to make it grammatically correct.
[0337] Step 5:
[0338] The server centralizes the translated assignments from all employees. The server stores the translated and grammar-checked text in a database and aggregates input from different employees.
[0339] Step 6:
[0340] The server passes the aggregated data to the generative AI to create a summary. The server extracts multiple issues stored in the database and passes them to the generative AI module to create a summary.
[0341] Step 7:
[0342] The server sends the generated summary to the administrator interface, where it displays the summary created by the generating AI in real time, allowing for immediate confirmation.
[0343] Step 8:
[0344] The administrator checks the summary on the management screen and makes a decision. The administrator evaluates the summary content through the interface and takes necessary action or makes a decision.
[0345] Through specific actions at each step, the entire system functions efficiently, supporting information sharing and quick decision-making among employees who speak different languages.
[0346] Example 1
[0347] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0348] In today's global companies, when employees who speak different languages work together, language barriers become an obstacle, making it difficult to communicate efficiently and make quick decisions. To address this issue, systems are needed that can efficiently translate information, check grammar, aggregate data, and summarize it. In existing systems, these processes are often performed manually, requiring time and effort. Furthermore, a lack of centralized information management makes it difficult for managers to quickly obtain the information they need.
[0349] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0350] In this invention, the server includes means for receiving information entered in different languages, means for translating the received information into a common language, means for checking and correcting the grammar of the translated information, means for centrally managing multiple pieces of information, means for summarizing the centrally managed information using a generative model, and means for reporting the summarized information to an administrator in real time. This enables smooth communication across language barriers and efficient aggregation and summarization of information. As a result, administrators can quickly obtain the information they need and make quick decisions.
[0351] "Means for receiving" refers to the function for importing information entered in different languages from a terminal into a server.
[0352] "Means for translating into a common language" refers to a function for converting received information from a different language into one pre-specified language (primarily English).
[0353] "Grammar checking and correction means" refers to functionality for detecting and, if necessary, correcting grammatical errors in the translated information.
[0354] "Means for centralized management" refers to a function for aggregating information received from multiple users into a central database and managing it in a unified manner.
[0355] "Means for summarizing using a generative model" refers to a function that uses a generative AI model based on aggregated information to extract important points and create a summary.
[0356] "Means of reporting in real time" refers to a function that allows the administrator to be notified of the generated summary immediately without any time lag.
[0357] MODE FOR CARRYING OUT THE INVENTION
[0358] This invention is a system for a multilingual collaborative task aggregation AI service. This system consists of three elements: a server, a terminal, and a user. Each element has a specific function and works together to solve business tasks.
[0359] Handling User Input
[0360] Users input business issues into the terminal in their native language. This input is done through a text field, and after the user describes the issue, they press the "Submit" button. This operation causes the terminal to generate data and send it to the server as JSON-formatted data. For example, if a user inputs "Customer feedback is delayed," this text is sent from the terminal to the server.
[0361] Translation Processing
[0362] The server receives the data sent from the device. It analyzes the received data and passes it to a multilingual translation AI module. This translation AI module may use, for example, the Google Translate API or the DeepL API. The input Japanese text is translated into English, the common language. For example, "Customer feedback is delayed" is translated to "Customer feedback is delayed."
[0363] Grammar Check
[0364] The translated text is sent to a grammar checker module, which can use the Grammarly API or LanguageTool API, for example. The grammar checker detects grammatical errors and corrects them if necessary. For example, "Customer feedback is delayed." is checked for grammar and any necessary corrections are made.
[0365] Data Aggregation
[0366] The server aggregates all the translated texts sent by users into a database. This process uses MySQL or PostgreSQL as a database management system (DBMS), which allows for centralized management of each user's input.
[0367] Data Summary
[0368] The aggregated data is passed to a generative AI module, which uses generative AI models such as OpenAI's GPT-3 and ChatGPT. The generative AI learns from past data and patterns, extracts key points, and summarizes the text. For example, if multiple users submit similar reports, the generative AI will generate a summary such as "Several employees report delays in customer feedback."
[0369] Reporting to management
[0370] The generated summary is sent to an administrator interface, where a dashboard is displayed that updates in real time, allowing administrators to instantly view the summarized information. For example, when an administrator opens the dashboard, they might see a summary that reads, "Several employees report delays in customer feedback."
[0371] Example prompt sentences
[0372] Below are some specific examples of prompt sentences to input to the generative AI model.
[0373] 1. The user types "Customer feedback is delayed" in Japanese into the device and presses the send button.
[0374] 2. The device sends this input to the server as JSON format data.
[0375] 3. The server passes the received data to the Google Translate API and translates it into English: "Customer feedback is delayed."
[0376] 4. Check the grammar of the translated text using the Grammarly API and make corrections if necessary.
[0377] 5. The server stores and aggregates all user data in a MySQL database.
[0378] 6. Pass the aggregated data to ChatGPT and generate a summary: “Several employees report delays in customer feedback.”
[0379] 7. The server sends a summary to the administrator's dashboard, where the administrator can instantly view it.
[0380] In this way, the present invention enables smooth communication in a multilingual environment, and supports efficient information gathering and summarization, as well as rapid decision-making.
[0381] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0382] Step 1:
[0383] The user inputs the issue into the terminal. The user writes the business issue in their native language into the text field on the terminal and presses the "Submit" button. This operation converts the entered text into JSON format by the terminal and sends it to the server. The input is the text "Customer feedback is delayed," and the output is JSON format data.
[0384] Step 2:
[0385] The server receives the data sent from the device. It analyzes the received JSON data and passes it to a multilingual translation AI module. This translation AI module uses, for example, a multilingual translation API. It receives JSON data as input and obtains translated text data as output. For example, the Japanese data "Customer feedback is delayed" is translated into the English data "Customer feedback is delayed."
[0386] Step 3:
[0387] The server sends the translated text to a grammar checker module. The grammar checker module uses, for example, a grammar check API. It receives the translated text data as input, detects grammatical errors, and outputs corrected text data as necessary. For example, the text "Customer feedback is delayed." is grammatically checked to ensure there are no errors.
[0388] Step 4:
[0389] The server aggregates all translated text submitted by users into a database. The database management system (DBMS) can be, for example, a relational database. It receives translated text data as input and outputs it as a centralized database record. For example, text submitted by multiple users can be aggregated into a single database.
[0390] Step 5:
[0391] The server passes the aggregated data to a generative AI module. This module uses a generative AI model that takes the aggregated data from the database as input, extracts key points, and outputs summarized text data. For example, multiple matching reports might be summarized as "Several employees report delays in customer feedback."
[0392] Step 6:
[0393] The server sends the generated summary to an administrator interface, which displays a dashboard that updates in real time, allowing administrators to instantly view the summarized information. The interface receives summarized text data as input and displays it as an administrator-accessible dashboard as output. For example, a real-time summary such as "Several employees report delays in customer feedback" appears on the administrator's dashboard.
[0394] (Application example 1)
[0395] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0396] In an environment where employees speak a wide variety of languages, it is necessary to quickly and efficiently collect, translate, and summarize business issues and provide accurate information to managers. Real-time problem reporting and rapid decision-making are particularly important in logistics centers. However, the overlap of different languages and manual information processing can lead to delays and errors in the collection, translation, and transmission of information.
[0397] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0398] In this invention, the server includes a means for receiving information input in different languages, a means for translating the received information into a common language, and a means for checking and correcting the grammar of the translated information. This provides a means for centrally managing multiple pieces of information, a means for summarizing the centrally managed information, and a means for reporting the summarized information to an administrator. This allows for the input of information in real time using a smart device, and for the input information to be translated, checked for grammar, and summarized, and then reported promptly to an administrator.
[0399] "Information entered in different languages" refers to text data entered by users in their various native languages.
[0400] A "common language" is a unified language (e.g., English) that is translated from different languages.
[0401] "Grammar checking and correction" is the process of making the translated text grammatically correct.
[0402] "Centralized management" refers to the process of aggregating multiple pieces of information into a central database and managing them in an integrated manner.
[0403] "Summarizing" refers to extracting important points from aggregated information and summarizing them concisely.
[0404] "Reporting to administrator" is the process of sending the processed information to an administrator interface so that the administrator can review it.
[0405] "Inputting information in real time using a smart device" refers to inputting information instantly from a mobile device such as a smartphone or tablet.
[0406] A "generative model" is an algorithm that uses machine learning techniques to summarize and generate information.
[0407] This paper describes a system that realizes a multilingual collaborative task aggregation AI service in a logistics center. This system consists of three main elements: a server, a terminal, and a user, each of which has a specific function.
[0408] Handling User Input
[0409] Users input their work-related tasks in their native language using a smartphone. This input is done through a text field, and after the user has written the task, they press the "Submit" button to complete the task. The device then sends the user's input as data to the server.
[0410] Translation and Grammar Check
[0411] The server receives user input sent from the device. The received data is first translated into a common language (e.g., English) using the Google Translator API. This translated text is then sent to the grammar checking module via OpenAI's API, where grammatical errors are checked and corrected.
[0412] Aggregation and Summarization
[0413] The server centrally manages translated assignments collected from multiple users and aggregates them into a database. The aggregated data is then extracted and summarized by OpenAI's generative AI model. This generative AI model learns from past data and patterns to efficiently and accurately summarize information.
[0414] Report to the administrator
[0415] The server sends the generated summary to an administrator interface, which updates in real time, allowing administrators to instantly see the summarized information and make quick decisions.
[0416] Specific examples
[0417] If a user inputs an issue in Japanese such as "Errors in the inventory management system occur frequently," the input is sent to the server via the terminal. The server translates this input into English using the Google Translator API, resulting in "Errors in the inventory management system occur frequently." The translated text is then checked for grammar using the OpenAI API, and corrected to "Errors in the inventory management system are occurring frequently." The generation AI generates a summary from many similar reports: "Frequent errors in inventory system." This summary is sent to the administrator interface, where the administrator can take prompt action based on it.
[0418] Example prompt sentence:
[0419] Grammar check prompt: "Please correct the following text: Errors in the inventory management system occur frequently."
[0420] Summary prompt: "Summarize the following text: Errors in the inventory management system are occurring frequently."
[0421] In this way, the present invention enables multilingual employees to efficiently gather information, translate, check grammar, aggregate data, and summarize, and quickly report to management, thereby supporting smooth communication and quick decision-making across language barriers in logistics centers.
[0422] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0423] Step 1:
[0424] A user uses a smart device to input a business issue in their native language. Specifically, the user writes the issue in a text field on the smartphone application and presses the "Submit" button. The input issue (e.g., "Errors are occurring frequently in the inventory management system") is sent from the device to the server. The input data is text data in the user's native language.
[0425] Step 2:
[0426] The server receives user input sent from the device. The received data is first sent to the Google Translator API, where the native text data is translated into a common language (e.g., English). The translated text data (e.g., "Errors in the inventory management system occur frequently.") is generated.
[0427] Step 3:
[0428] The server sends the translated text data to OpenAI's API, where a grammar check is performed using a generative AI model. Specifically, the server detects and corrects grammatical errors in the translated text, generating corrected text data (e.g., "Errors in the inventory management system are occurring frequently.").
[0429] Step 4:
[0430] The server stores the grammar-checked text data in a database and manages it centrally. Multiple similar data sets are aggregated, and the text data is stored in a management database. This ensures that assignments from multiple users are managed in an integrated manner.
[0431] Step 5:
[0432] The server sends the aggregated data to OpenAI's generative AI model, which extracts key points and creates a summary. Specifically, the server selects key content from multiple translated and corrected text data and generates a summary text (e.g., "Frequent errors in inventory system.").
[0433] Step 6:
[0434] The server sends the generated summary text to an administrator interface, where administrators can view the summarized information using an interface that updates in real time, providing administrators with accurate information for quick decision making.
[0435] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0436] The present invention is a system that incorporates an emotion engine that recognizes the user's emotions, as well as a means for receiving information entered in different languages, translating it into a common language, and checking and correcting the grammar of the information. This system consists of three main components: a server, a terminal, and a user, each of which has a specific function.
[0437] User Input Processing and Emotion Recognition
[0438] Users input work-related issues and information into the device in their native language. After selecting the language they want to use and entering the issue in the text field, they press the "Submit" button. The device then sends this input data to the server. During this process, the device also collects emotional information along with the user's input data. The emotional information is recognized by an emotion engine that analyzes the user's input content and behavior during input.
[0439] Translation and Grammar Check
[0440] The server receives user input and emotion information sent from the device. The received data is first passed to a multilingual translation AI module, which translates the input language into a common language (e.g., English). The translated text is then sent to a grammar checker module, where it is checked for grammatical errors and corrected if necessary.
[0441] Unified management of emotional information
[0442] The server centrally manages the emotional information along with the translated and grammar-checked information. The server stores this information in a database and aggregates input data from different employees. At this time, the emotional information is treated like regular text data and stored in the database.
[0443] Aggregation and Summarization
[0444] The server aggregates multiple issues stored in the database and passes them to the generation AI to create a summary. Emotional information is also provided to the generation AI, and the tendency and intensity of emotions are reflected in the summarization process. For example, if many employees are feeling stressed, that emotional information will be noted as part of the summary.
[0445] Report to the administrator
[0446] The server sends the generated summary and emotion information to the administrator interface, which is updated in real time, allowing the administrator to instantly check the summary and emotion information. This allows the administrator to understand not only the text information but also the emotional state of employees, enabling them to make better decisions.
[0447] Specific examples
[0448] If a user inputs an issue in Japanese, such as "Customer feedback is delayed. I am very upset," and the input includes emotional expressions, the input is sent to the server via the terminal. The server translates this into English as "Customer feedback is delayed. I am very upset." A grammar checker then corrects the input to make it grammatically correct. The server also uses an emotion engine to recognize emotions such as "confused" and "stressed." Next, multiple similar reports are aggregated, and a generative AI creates a summary: "Several employees report delays in customer feedback and are very stressed." Finally, this summary and emotional information are sent to the administrator interface, where the administrator reviews it and considers countermeasures.
[0449] This system integrates multilingual information collection, translation, grammar checking, and emotion recognition, and can centrally manage and summarize data for quick reporting, eliminating communication barriers caused by language and cultural differences and supporting appropriate decision-making that takes into account the emotional state of employees.
[0450] The processing flow will be explained below.
[0451] Step 1:
[0452] The user enters work-related issues and information into the device in their native language. After selecting the language they want to use and entering the issue in the text field, they press the "Submit" button. This operation also collects emotional information.
[0453] Step 2:
[0454] The device sends the user's input and emotional information to the server. The input content and emotional information are sent to the server as an HTTP request. The emotional information is the result of an emotion engine analyzing the user's input content and behavior at the time of input.
[0455] Step 3:
[0456] The server passes the received information to a multilingual translation AI, which translates it into a common language (e.g., English). The server analyzes the received data, sends the text and language information to the translation API, and obtains the translation results.
[0457] Step 4:
[0458] The server passes the translated text to the grammar checker module, which checks and corrects grammatical errors, and the translation result is sent to the grammar checker API, which formats it to be grammatically correct.
[0459] Step 5:
[0460] The server centralizes the translated and grammar-checked information, as well as the sentiment information, which is then stored in a database that aggregates input from different employees.
[0461] Step 6:
[0462] The server passes the collected data to the generation AI to create a summary. Using the issue information and emotion information stored in the database, the generation AI module extracts important points and reflects them in the summary.
[0463] Step 7:
[0464] The server sends the generated summary and emotion information to the administrator interface, which is updated in real time, allowing the administrator to instantly check the summary and emotion information.
[0465] Step 8:
[0466] The administrator checks the summary on the management screen and makes a decision. The administrator evaluates the summary content and sentiment information through the interface and quickly takes necessary actions and makes decisions.
[0467] As a specific example, if a user inputs and submits the issue in Japanese, "Customer feedback is delayed. I am very upset.", the input and emotions are recognized as "confused" and "stressed." The server translates the received input into English, resulting in "Customer feedback is delayed. I am very upset." The grammar checker corrects the input to make it grammatically correct, and then stores it in the database. Next, it is aggregated with similar input from other employees, and the generative AI creates a summary: "Several employees report delays in customer feedback and are very stressed." This summary and emotional information are sent to the administrator interface, where administrators can make quick decisions based on it.
[0468] In this way, the entire system functions efficiently, supporting information sharing among employees who speak different languages and quick decision-making that takes emotional information into account.
[0469] Example 2
[0470] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0471] There is a need for a system that can efficiently collect, translate, check grammar, and analyze emotions from users who speak different languages, and then report the results to administrators. This system must also be able to accurately analyze users' emotional information and generate summaries that take this information into account, thereby solving the problem of providing administrators with information that includes the users' emotional state.
[0472] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0473] In this invention, the server includes means for receiving information input in different languages, means for translating the received information into a common language, means for checking and correcting the grammar of the translated information, means for centrally managing multiple pieces of information, means for summarizing the centrally managed information, means for reporting the summarized information to an administrator, means for collecting emotional information from the input information, means for analyzing the emotional information, and means for generating a summary that includes the emotional information. This makes it possible to efficiently collect and manage information from users who speak different languages and to provide information that includes their emotional states.
[0474] "Different languages" refers to languages used by users in multiple different countries or regions.
[0475] "Input information" refers to text data and assignments entered by the user using the terminal.
[0476] "Means of receiving" refers to the mechanism or method by which the server receives data sent from the terminal.
[0477] "Common language" refers to a standard language that is used uniformly within the system, such as English.
[0478] "Means of translation" refers to the functions and methods for converting information entered in different languages into a common language.
[0479] "Grammar checking and correction means" refers to a method or module for detecting grammatical errors in the translated text and correcting it to make it grammatically correct.
[0480] "Means of centralized management" refers to a method of integrating multiple input data and translation results and centrally managing them in a database or management system.
[0481] "Means of summarization" refers to techniques or methods, such as the use of generative AI models, to succinctly summarize large amounts of input information or data.
[0482] "Means for reporting to administrator" refers to the method or interface for notifying the administrator of the summarized information.
[0483] "Emotional information" refers to emotional states analyzed from information entered by the user, such as data on "confusion" or "stress."
[0484] "Means for collecting emotional information" refers to the technology or method for analyzing user input data to recognize and collect emotions.
[0485] "Means for analyzing emotional information" refers to technology for analyzing collected emotional data and evaluating its content and trends.
[0486] "Means for generating summaries that include emotional information" refers to techniques and methods for creating summaries of input data that take emotional information into account.
[0487] The present invention is a system that receives information entered in different languages, translates it into a common language, and checks and corrects the grammar of the information. This system also has the function of collecting and analyzing emotional information, centrally managing the collected information, and finally summarizing and reporting it to an administrator. Specific embodiments of this system are described below.
[0488] Users input business issues and information into the device in their native language. For example, they can enter something like "Feedback from customers is delayed. We are having a lot of trouble" in Japanese and press the "Send" button. Users can first select the language they want to use and then enter the issue or information in the text field.
[0489] The device receives the user's input data. At the same time, a built-in emotion engine simultaneously collects emotional information based on the input content and the user's input speed. This emotion engine analyzes and recognizes emotions such as "confusion" or "stress" from the user's text content.
[0490] The server receives the user's input data and emotional information sent from the device. The received data is first passed to a multilingual translation AI module, which translates the different input languages (such as Japanese) into a common language (such as English). An industry-standard AI translation engine is often used as the translation module.
[0491] The translated text is then sent to a grammar checker module, which can incorporate existing grammar checking software such as Grammarly, to check for grammatical errors and correct them if necessary.
[0492] Once the translation and grammar check are complete, the information is centrally managed by a server. This information is stored in a database that aggregates input data from different employees. This database management system may use MySQL or PostgreSQL.
[0493] Next, the server aggregates multiple issues stored in the database and passes them to a generation AI to create a summary. Here, emotional information is also provided to the generation AI, and the tendency and intensity of emotions are reflected in the summarization process. For example, a summary is generated that includes emotional information, such as "Most employees are stressed by delays in feedback from customers."
[0494] Finally, the server sends the generated summary and emotion information to the administrator interface, which is updated in real time, allowing the administrator to instantly check the information. This allows the administrator to understand the employee's emotional state and provides information for better decision-making.
[0495] For example, if a user inputs an issue in Japanese such as "Customer feedback is delayed. I am very upset." and the input contains emotional expressions, the input is sent to the server via the terminal. The server translates this into English as "Customer feedback is delayed. I am very upset." A grammar checker then corrects the input to make it grammatically correct. The server also uses an emotion engine to recognize emotions such as "confused" and "stressed." Next, multiple similar reports are aggregated, and a generative AI creates a summary such as "Several employees report delays in customer feedback and are very stressed." Finally, this summary and emotional information are sent to the administrator interface, where the administrator reviews it and considers countermeasures.
[0496] An example prompt is, "Recognize the sentiment from the user's input data and summarize the text in a way that includes sentiment. Summarize the text as follows: 'Several employees report delays in customer feedback and are very stressed.'"
[0497] As described above, the present invention can provide a system that can efficiently manage information regardless of language differences and can report information including emotional information to a manager.
[0498] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0499] Step 1:
[0500] The user inputs the issue or information into the device in their native language. After inputting, they select the language they want to use, enter something like "Feedback from customers is delayed. We are having a lot of trouble," in the text field, and press the "Send" button. This saves the input information as text data on the device.
[0501] Step 2:
[0502] The device receives text data entered by the user. At the same time, the emotion engine analyzes the text content and behavioral data such as the user's typing speed, and collects and recognizes emotional information such as "confusion" or "stress." The input is text data, and the output is text data and emotional information. Specifically, the device also collects the user's keystroke data and sends it to the emotion engine for analysis.
[0503] Step 3:
[0504] The device sends the user's text data and emotional information to the server. Specifically, the device requests the server to send data using the HTTP protocol. The input is text data and emotional information, and the output is sent to the server.
[0505] Step 4:
[0506] The server receives text data and emotion information sent from the device. The received data is first passed to a multilingual translation AI module, which translates the input language (e.g., Japanese) into a common language (e.g., English). The input is Japanese text data, and the output is English text data. The server calls the translation engine to convert the data.
[0507] Step 5:
[0508] The server sends the translated text data to the grammar checker module, which checks and corrects grammatical errors. The input is English text data, and the output is English text data with corrected grammar. The server passes the data to the grammar checker module and receives the corrected results.
[0509] Step 6:
[0510] The server stores the corrected text data and emotion information in a database for centralized management. The input is the corrected text data and emotion information, and the output is saving to the database. The server inserts data into the database management system using SQL queries.
[0511] Step 7:
[0512] The server aggregates multiple issues and emotion information stored in the database and passes it to the generative AI model to create a summary. The input is multiple data in the database, and the output is summary data. Specifically, the server retrieves the data using an SQL query and provides it to the generative AI model along with a prompt statement to generate a summary. An example of a prompt statement is "Several employees report delays in customer feedback and are very stressed."
[0513] Step 8:
[0514] The server sends the generated summary and emotion information to the administrator interface. The input is summary data and emotion information, and the output is interface updates. The server sends data to the administrator's browser in real time using WebSocket or HTTP protocol and updates the interface.
[0515] Through these steps, the system can efficiently translate, check grammar, analyze sentiment, and aggregate information in different languages to generate summaries and provide them to administrators.
[0516] (Application example 2)
[0517] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0518] In the past, when communicating between different languages, language translation, grammar checking, and emotion recognition tended to be performed in separate systems, making it difficult to centrally manage and summarize information and report it to managers immediately.In addition, in physical stores, there was a lack of means to receive feedback from multilingual customers in real time and process it smoothly, which placed a heavy burden on staff.
[0519] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving information input in different languages, means for translating the received information into a common language, means for checking and correcting the grammar of the translated information, means for recognizing the user's emotions, means for centrally managing multiple pieces of information, means for summarizing the centrally managed information, means for reporting the summarized information to a manager, and components that allow store staff to input customer feedback on-site in real time and instantly check the translation results and emotional information. This centralizes feedback processing in a multilingual environment, making it possible to simultaneously manage customer feedback and the emotional state of staff.
[0520] The "means for receiving information input in a different language" is a function for receiving information input by a user in a different language at a terminal.
[0521] The "means for translating received information into a common language" is a function for automatically translating received information in multiple languages into a common language.
[0522] The "means for checking and correcting the grammar of translated information" is a function for detecting grammatical errors in translated information and automatically correcting them.
[0523] The "means for recognizing the user's emotions" is a function that analyzes the user's input content and behavior at the time of input and identifies the user's emotional state.
[0524] "Means for centrally managing multiple pieces of information" refers to a function that consolidates and manages multiple pieces of input information and related emotional information in one place.
[0525] "Means for summarizing centrally managed information" is a function that generates a short summary of the main points based on the aggregated information.
[0526] The "means for reporting summarized information to the administrator" is a function for instantly notifying the administrator of summarized information and emotional information.
[0527] "A component that allows store staff to input customer feedback in real time on-site and instantly check the translation results and emotional information" is a device that allows store staff to input feedback from customers in real time and instantly view it along with the translation and emotional analysis results.
[0528] This invention is a unified management system that receives customer feedback in real time in a multilingual environment in a brick-and-mortar store, translates it, checks grammar, and recognizes emotions. Specific embodiments of this system are described below.
[0529] System configuration
[0530] The system mainly consists of three main components: a server, a terminal, and a user.
[0531] Server: Translates data, recognizes emotions, checks grammar, and centralizes and summarizes information.
[0532] Terminal: A device where store clerks can input customer feedback in real time and check the translated results and sentiment analysis results. This can be a smartphone, smart glasses, or a head-mounted display.
[0533] Users: Includes customers and store clerks, where customers provide feedback and store clerks input it into the terminal.
[0534] Program Overview
[0535] The server receives information (feedback) entered in different languages and translates it into a common language. Translation services such as the Google Cloud Translation API are used for translation. The translated text is then checked for grammar using services such as the Grammarly API and corrected as necessary. Emotion recognition is performed using services such as the Microsoft Azure Emotion API to analyze the user's emotional state.
[0536] The received information and analyzed sentiment information are centrally managed on a server. Multiple feedback information is stored in a database and summarized using a generative AI model, allowing managers to efficiently grasp the key points.
[0537] Hardware and software used
[0538] Hardware:
[0539] Smartphone
[0540] Smart Glasses
[0541] head-mounted display
[0542] software:
[0543] Google Cloud Translation API (translation)
[0544] Microsoft Azure Emotion API (emotion recognition)
[0545] Grammarly API (Grammar Check)
[0546] Processing Flow
[0547] 1. User (store clerk): Enters customer feedback into a device such as a smartphone.
[0548] 2. Terminal: Sends the entered information to the server.
[0549] 3. Server:
[0550] Translate information.
[0551] Check and correct the grammar of the translated text.
[0552] Recognize user emotions.
[0553] 4. Server: Centralizes the translated and analyzed information, generates summaries, and reports to the administrator.
[0554] 5. Terminal: Displays translation results and sentiment analysis results in real time.
[0555] Specific examples
[0556] A customer provides feedback in Japanese, saying, "Customer feedback is delayed. I am very upset." The store clerk enters this into a terminal. This input is sent to the server, where it is translated into English, becoming, "Customer feedback is delayed. I am very upset." Grammar checks are then performed to ensure the correct form is achieved. Emotion recognition identifies emotions such as "confusion" and "stress," and the aggregated data generates a summary: "Several employees report delays in customer feedback and are very stressed." Finally, this summary and emotional information are reported to management.
[0557] Prompt Sentence Examples
[0558] "Please translate the feedback below, check for grammar, and recognize emotional information.
[0559] Language: Japanese
[0560] Feedback: "Customers are very unhappy"
[0561] As described above, it is possible to simultaneously manage customer feedback and the emotional state of staff in a multilingual environment, and a system is provided that can instantly translate and respond to communication between store staff and customers.
[0562] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0563] Step 1:
[0564] User (store clerk): The store clerk uses a device such as a smartphone or smart glasses to input customer feedback in real time. The input data is in text format, for example, "Customer feedback is delayed. We are very troubled." The input data is temporarily stored in the device.
[0565] Step 2:
[0566] Terminal: The terminal sends the input text data to the server, along with input language information (e.g., feedback in Japanese and its timestamp). A secure communication protocol (e.g., HTTPS) is used for this data transmission.
[0567] Step 3:
[0568] Server: The server receives the text data and first translates it into a common language (e.g., English) using the Google Cloud Translation API. The input is the original Japanese feedback, and the output is the translated data, such as "Customer feedback is delayed. I am very upset."
[0569] Step 4:
[0570] Server: After getting the translated data, we use the Grammarly API to perform a grammar check. Here, the input is the translated English text and the output is the grammatically corrected text. The corrections could be, for example, "Customer feedback has been delayed. I am very upset."
[0571] Step 5:
[0572] Server: After the grammar check is complete, emotion recognition is performed using the Microsoft Azure Emotion API. The input is the user's text data ("Customer feedback has been delayed. I am very upset."), and the output is the recognized emotion information (e.g., "confused" or "stressed").
[0573] Step 6:
[0574] Server: The data that has been translated, checked for grammar, and recognized for emotion is managed in a centralized database. The server stores this data and also assigns metadata such as timestamps to each piece of data. The input is the original feedback and various analyzed data, and the output is a centralized database entry.
[0575] Step 7:
[0576] Server: Multiple pieces of feedback stored in a database are summarized using a generative AI model. The input is a set of accumulated database entries, and the output is summary information, such as "Several employees report delays in customer feedback and are very stressed."
[0577] Step 8:
[0578] Server: Finally, the generated summary information and related sentiment information are sent to the administrator interface, allowing the administrator to check information based on staff status and customer feedback in real time. The input is the summarized data, and the output is the display results on the interface accessed by the administrator.
[0579] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0580] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0581] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0582] [Third embodiment]
[0583] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0584] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0585] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0586] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0587] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0588] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0589] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0590] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0591] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0592] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0593] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0594] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0595] As a concrete example of the present invention, we will explain a multilingual collaborative task aggregation AI service. This system consists of three main components: a server, a terminal, and a user, each of which has a specific function.
[0596] Handling User Input
[0597] Users input their business issues into the terminal in their native language. This input is done through a text field, and after the user has written the issue, they press the "Submit" button. This operation causes the terminal to send the user's input as data to the server.
[0598] Translation and Grammar Check
[0599] The server receives user input sent from the terminal. The received data is first passed to a multilingual translation AI module, which translates the input language into a common language (e.g., English). The translated text is then sent to a grammar checker module, where it is checked for grammatical errors and corrected if necessary.
[0600] Aggregation and Summarization
[0601] The server centralizes translated assignments collected from multiple users. This process involves aggregating each user's input into a database. The aggregated data is then passed to a generative AI module, which further extracts and summarizes key points. This generative AI module has learned from past data and patterns to efficiently and accurately summarize information.
[0602] Reporting to management
[0603] The server sends the generated summary to an administrator interface, which updates in real time, allowing the administrator to instantly view the summarized information and make quick decisions based on the summary.
[0604] Specific examples
[0605] If a user inputs the issue "Customer feedback is delayed" in Japanese, the input is sent to the server via the terminal. The server translates this input into English as "Customer feedback is delayed." The grammar checker then corrects it to make it grammatically correct. Next, multiple similar reports are aggregated and a generative AI creates a summary: "Several employees report delays in customer feedback." Finally, this summary is sent to the administrator interface, where it is reviewed and evaluated by the administrator.
[0606] In this way, the present invention helps multilingual employees efficiently gather information, translate, check grammar, aggregate data, and summarize it, and quickly report it to managers, thereby supporting smooth communication and quick decision-making across language barriers.
[0607] The processing flow will be explained below.
[0608] Step 1:
[0609] The user enters business issues and information into the terminal in their native language. After selecting the language they want to use and entering the issue in the text field, they press the "Submit" button.
[0610] Step 2:
[0611] The terminal sends the user's input to the server. The terminal sends the input data to the server as an HTTP request. This data includes the language information selected by the user and the assignment text.
[0612] Step 3:
[0613] The server passes the received information to a multilingual translation AI, which translates it into a common language (e.g., English). The server analyzes the received data, sends the text and language information to the translation API, and obtains the translation results.
[0614] Step 4:
[0615] The server passes the translated text to the grammar checker module to check and correct grammatical errors. The server sends the translated text to the grammar checker API to make it grammatically correct.
[0616] Step 5:
[0617] The server centralizes the translated assignments from all employees. The server stores the translated and grammar-checked text in a database and aggregates input from different employees.
[0618] Step 6:
[0619] The server passes the aggregated data to the generative AI to create a summary. The server extracts multiple issues stored in the database and passes them to the generative AI module to create a summary.
[0620] Step 7:
[0621] The server sends the generated summary to the administrator interface, where it displays the summary created by the generating AI in real time, allowing for immediate confirmation.
[0622] Step 8:
[0623] The administrator checks the summary on the management screen and makes a decision. The administrator evaluates the summary content through the interface and takes necessary action or makes a decision.
[0624] Through specific actions at each step, the entire system functions efficiently, supporting information sharing and quick decision-making among employees who speak different languages.
[0625] Example 1
[0626] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0627] In today's global companies, when employees who speak different languages work together, language barriers become an obstacle, making it difficult to communicate efficiently and make quick decisions. To address this issue, systems are needed that can efficiently translate information, check grammar, aggregate data, and summarize it. In existing systems, these processes are often performed manually, requiring time and effort. Furthermore, a lack of centralized information management makes it difficult for managers to quickly obtain the information they need.
[0628] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0629] In this invention, the server includes means for receiving information entered in different languages, means for translating the received information into a common language, means for checking and correcting the grammar of the translated information, means for centrally managing multiple pieces of information, means for summarizing the centrally managed information using a generative model, and means for reporting the summarized information to an administrator in real time. This enables smooth communication across language barriers and efficient aggregation and summarization of information. As a result, administrators can quickly obtain the information they need and make quick decisions.
[0630] "Means for receiving" refers to the function for importing information entered in different languages from a terminal into a server.
[0631] "Means for translating into a common language" refers to a function for converting received information from a different language into one pre-specified language (primarily English).
[0632] "Grammar checking and correction means" refers to functionality for detecting and, if necessary, correcting grammatical errors in the translated information.
[0633] "Means for centralized management" refers to a function for aggregating information received from multiple users into a central database and managing it in a unified manner.
[0634] "Means for summarizing using a generative model" refers to a function that uses a generative AI model based on aggregated information to extract important points and create a summary.
[0635] "Means of reporting in real time" refers to a function that allows the administrator to be notified of the generated summary immediately without any time lag.
[0636] MODE FOR CARRYING OUT THE INVENTION
[0637] This invention is a system for a multilingual collaborative task aggregation AI service. This system consists of three elements: a server, a terminal, and a user. Each element has a specific function and works together to solve business tasks.
[0638] Handling User Input
[0639] Users input business issues into the terminal in their native language. This input is done through a text field, and after the user describes the issue, they press the "Submit" button. This operation causes the terminal to generate data and send it to the server as JSON-formatted data. For example, if a user inputs "Customer feedback is delayed," this text is sent from the terminal to the server.
[0640] Translation Processing
[0641] The server receives the data sent from the device. It analyzes the received data and passes it to a multilingual translation AI module. This translation AI module may use, for example, the Google Translate API or the DeepL API. The input Japanese text is translated into English, the common language. For example, "Customer feedback is delayed" is translated to "Customer feedback is delayed."
[0642] Grammar Check
[0643] The translated text is sent to a grammar checker module, which can use the Grammarly API or LanguageTool API, for example. The grammar checker detects grammatical errors and corrects them if necessary. For example, "Customer feedback is delayed." is checked for grammar and any necessary corrections are made.
[0644] Data Aggregation
[0645] The server aggregates all the translated texts sent by users into a database. This process uses MySQL or PostgreSQL as a database management system (DBMS), which allows for centralized management of each user's input.
[0646] Data Summary
[0647] The aggregated data is passed to a generative AI module, which uses generative AI models such as OpenAI's GPT-3 and ChatGPT. The generative AI learns from past data and patterns, extracts key points, and summarizes the text. For example, if multiple users submit similar reports, the generative AI will generate a summary such as "Several employees report delays in customer feedback."
[0648] Reporting to management
[0649] The generated summary is sent to an administrator interface, where a dashboard is displayed that updates in real time, allowing administrators to instantly view the summarized information. For example, when an administrator opens the dashboard, they might see a summary that reads, "Several employees report delays in customer feedback."
[0650] Example prompt sentences
[0651] Below are some specific examples of prompt sentences to input to the generative AI model.
[0652] 1. The user types "Customer feedback is delayed" in Japanese into the device and presses the send button.
[0653] 2. The device sends this input to the server as JSON format data.
[0654] 3. The server passes the received data to the Google Translate API and translates it into English: "Customer feedback is delayed."
[0655] 4. Check the grammar of the translated text using the Grammarly API and make corrections if necessary.
[0656] 5. The server stores and aggregates all user data in a MySQL database.
[0657] 6. Pass the aggregated data to ChatGPT and generate a summary: “Several employees report delays in customer feedback.”
[0658] 7. The server sends a summary to the administrator's dashboard, where the administrator can instantly view it.
[0659] In this way, the present invention enables smooth communication in a multilingual environment, and supports efficient information gathering and summarization, as well as rapid decision-making.
[0660] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0661] Step 1:
[0662] The user inputs the issue into the terminal. The user writes the business issue in their native language into the text field on the terminal and presses the "Submit" button. This operation converts the entered text into JSON format by the terminal and sends it to the server. The input is the text "Customer feedback is delayed," and the output is JSON format data.
[0663] Step 2:
[0664] The server receives the data sent from the device. It analyzes the received JSON data and passes it to a multilingual translation AI module. This translation AI module uses, for example, a multilingual translation API. It receives JSON data as input and obtains translated text data as output. For example, the Japanese data "Customer feedback is delayed" is translated into the English data "Customer feedback is delayed."
[0665] Step 3:
[0666] The server sends the translated text to a grammar checker module. The grammar checker module uses, for example, a grammar check API. It receives the translated text data as input, detects grammatical errors, and outputs corrected text data as necessary. For example, the text "Customer feedback is delayed." is grammatically checked to ensure there are no errors.
[0667] Step 4:
[0668] The server aggregates all translated text submitted by users into a database. The database management system (DBMS) can be, for example, a relational database. It receives translated text data as input and outputs it as a centralized database record. For example, text submitted by multiple users can be aggregated into a single database.
[0669] Step 5:
[0670] The server passes the aggregated data to a generative AI module. This module uses a generative AI model that takes the aggregated data from the database as input, extracts key points, and outputs summarized text data. For example, multiple matching reports might be summarized as "Several employees report delays in customer feedback."
[0671] Step 6:
[0672] The server sends the generated summary to an administrator interface, which displays a dashboard that updates in real time, allowing administrators to instantly view the summarized information. The interface receives summarized text data as input and displays it as an administrator-accessible dashboard as output. For example, a real-time summary such as "Several employees report delays in customer feedback" appears on the administrator's dashboard.
[0673] (Application example 1)
[0674] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0675] In an environment where employees speak a wide variety of languages, it is necessary to quickly and efficiently collect, translate, and summarize business issues and provide accurate information to managers. Real-time problem reporting and rapid decision-making are particularly important in logistics centers. However, the overlap of different languages and manual information processing can lead to delays and errors in the collection, translation, and transmission of information.
[0676] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0677] In this invention, the server includes a means for receiving information input in different languages, a means for translating the received information into a common language, and a means for checking and correcting the grammar of the translated information. This provides a means for centrally managing multiple pieces of information, a means for summarizing the centrally managed information, and a means for reporting the summarized information to an administrator. This allows for the input of information in real time using a smart device, and for the input information to be translated, checked for grammar, and summarized, and then reported promptly to an administrator.
[0678] "Information entered in different languages" refers to text data entered by users in their various native languages.
[0679] A "common language" is a unified language (e.g., English) that is translated from different languages.
[0680] "Grammar checking and correction" is the process of making the translated text grammatically correct.
[0681] "Centralized management" refers to the process of aggregating multiple pieces of information into a central database and managing them in an integrated manner.
[0682] "Summarizing" refers to extracting important points from aggregated information and summarizing them concisely.
[0683] "Reporting to administrator" is the process of sending the processed information to an administrator interface so that the administrator can review it.
[0684] "Inputting information in real time using a smart device" refers to inputting information instantly from a mobile device such as a smartphone or tablet.
[0685] A "generative model" is an algorithm that uses machine learning techniques to summarize and generate information.
[0686] This paper describes a system that realizes a multilingual collaborative task aggregation AI service in a logistics center. This system consists of three main elements: a server, a terminal, and a user, each of which has a specific function.
[0687] Handling User Input
[0688] Users input their work-related tasks in their native language using a smartphone. This input is done through a text field, and after the user has written the task, they press the "Submit" button to complete the task. The device then sends the user's input as data to the server.
[0689] Translation and Grammar Check
[0690] The server receives user input sent from the device. The received data is first translated into a common language (e.g., English) using the Google Translator API. This translated text is then sent to the grammar checking module via OpenAI's API, where grammatical errors are checked and corrected.
[0691] Aggregation and Summarization
[0692] The server centrally manages translated assignments collected from multiple users and aggregates them into a database. The aggregated data is then extracted and summarized by OpenAI's generative AI model. This generative AI model learns from past data and patterns to efficiently and accurately summarize information.
[0693] Report to the administrator
[0694] The server sends the generated summary to an administrator interface, which updates in real time, allowing administrators to instantly see the summarized information and make quick decisions.
[0695] Specific examples
[0696] If a user inputs an issue in Japanese such as "Errors in the inventory management system occur frequently," the input is sent to the server via the terminal. The server translates this input into English using the Google Translator API, resulting in "Errors in the inventory management system occur frequently." The translated text is then checked for grammar using the OpenAI API, and corrected to "Errors in the inventory management system are occurring frequently." The generation AI generates a summary from many similar reports: "Frequent errors in inventory system." This summary is sent to the administrator interface, where the administrator can take prompt action based on it.
[0697] Example prompt sentence:
[0698] Grammar check prompt: "Please correct the following text: Errors in the inventory management system occur frequently."
[0699] Summary prompt: "Summarize the following text: Errors in the inventory management system are occurring frequently."
[0700] In this way, the present invention enables multilingual employees to efficiently gather information, translate, check grammar, aggregate data, and summarize, and quickly report to management, thereby supporting smooth communication and quick decision-making across language barriers in logistics centers.
[0701] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0702] Step 1:
[0703] A user uses a smart device to input a business issue in their native language. Specifically, the user writes the issue in a text field on the smartphone application and presses the "Submit" button. The input issue (e.g., "Errors are occurring frequently in the inventory management system") is sent from the device to the server. The input data is text data in the user's native language.
[0704] Step 2:
[0705] The server receives user input sent from the device. The received data is first sent to the Google Translator API, where the native text data is translated into a common language (e.g., English). The translated text data (e.g., "Errors in the inventory management system occur frequently.") is generated.
[0706] Step 3:
[0707] The server sends the translated text data to OpenAI's API, where a grammar check is performed using a generative AI model. Specifically, the server detects and corrects grammatical errors in the translated text, generating corrected text data (e.g., "Errors in the inventory management system are occurring frequently.").
[0708] Step 4:
[0709] The server stores the grammar-checked text data in a database and manages it centrally. Multiple similar data sets are aggregated, and the text data is stored in a management database. This ensures that assignments from multiple users are managed in an integrated manner.
[0710] Step 5:
[0711] The server sends the aggregated data to OpenAI's generative AI model, which extracts key points and creates a summary. Specifically, the server selects key content from multiple translated and corrected text data and generates a summary text (e.g., "Frequent errors in inventory system.").
[0712] Step 6:
[0713] The server sends the generated summary text to an administrator interface, where administrators can view the summarized information using an interface that updates in real time, providing administrators with accurate information for quick decision making.
[0714] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0715] The present invention is a system that incorporates an emotion engine that recognizes the user's emotions, as well as a means for receiving information entered in different languages, translating it into a common language, and checking and correcting the grammar of the information. This system consists of three main components: a server, a terminal, and a user, each of which has a specific function.
[0716] User Input Processing and Emotion Recognition
[0717] Users input work-related issues and information into the device in their native language. After selecting the language they want to use and entering the issue in the text field, they press the "Submit" button. The device then sends this input data to the server. During this process, the device also collects emotional information along with the user's input data. The emotional information is recognized by an emotion engine that analyzes the user's input content and behavior during input.
[0718] Translation and Grammar Check
[0719] The server receives user input and emotion information sent from the device. The received data is first passed to a multilingual translation AI module, which translates the input language into a common language (e.g., English). The translated text is then sent to a grammar checker module, where it is checked for grammatical errors and corrected if necessary.
[0720] Unified management of emotional information
[0721] The server centrally manages the emotional information along with the translated and grammar-checked information. The server stores this information in a database and aggregates input data from different employees. At this time, the emotional information is treated like regular text data and stored in the database.
[0722] Aggregation and Summarization
[0723] The server aggregates multiple issues stored in the database and passes them to the generation AI to create a summary. Emotional information is also provided to the generation AI, and the tendency and intensity of emotions are reflected in the summarization process. For example, if many employees are feeling stressed, that emotional information will be noted as part of the summary.
[0724] Report to the administrator
[0725] The server sends the generated summary and emotion information to the administrator interface, which is updated in real time, allowing the administrator to instantly check the summary and emotion information. This allows the administrator to understand not only the text information but also the emotional state of employees, enabling them to make better decisions.
[0726] Specific examples
[0727] If a user inputs an issue in Japanese, such as "Customer feedback is delayed. I am very upset," and the input includes emotional expressions, the input is sent to the server via the terminal. The server translates this into English as "Customer feedback is delayed. I am very upset." A grammar checker then corrects the input to make it grammatically correct. The server also uses an emotion engine to recognize emotions such as "confused" and "stressed." Next, multiple similar reports are aggregated, and a generative AI creates a summary: "Several employees report delays in customer feedback and are very stressed." Finally, this summary and emotional information are sent to the administrator interface, where the administrator reviews it and considers countermeasures.
[0728] This system integrates multilingual information collection, translation, grammar checking, and emotion recognition, and can centrally manage and summarize data for quick reporting, eliminating communication barriers caused by language and cultural differences and supporting appropriate decision-making that takes into account the emotional state of employees.
[0729] The processing flow will be explained below.
[0730] Step 1:
[0731] The user enters work-related issues and information into the device in their native language. After selecting the language they want to use and entering the issue in the text field, they press the "Submit" button. This operation also collects emotional information.
[0732] Step 2:
[0733] The device sends the user's input and emotional information to the server. The input content and emotional information are sent to the server as an HTTP request. The emotional information is the result of an emotion engine analyzing the user's input content and behavior at the time of input.
[0734] Step 3:
[0735] The server passes the received information to a multilingual translation AI, which translates it into a common language (e.g., English). The server analyzes the received data, sends the text and language information to the translation API, and obtains the translation results.
[0736] Step 4:
[0737] The server passes the translated text to the grammar checker module, which checks and corrects grammatical errors, and the translation result is sent to the grammar checker API, which formats it to be grammatically correct.
[0738] Step 5:
[0739] The server centralizes the translated and grammar-checked information, as well as the sentiment information, which is then stored in a database that aggregates input from different employees.
[0740] Step 6:
[0741] The server passes the collected data to the generation AI to create a summary. Using the issue information and emotion information stored in the database, the generation AI module extracts important points and reflects them in the summary.
[0742] Step 7:
[0743] The server sends the generated summary and emotion information to the administrator interface, which is updated in real time, allowing the administrator to instantly check the summary and emotion information.
[0744] Step 8:
[0745] The administrator checks the summary on the management screen and makes a decision. The administrator evaluates the summary content and sentiment information through the interface and quickly takes necessary actions and makes decisions.
[0746] As a specific example, if a user inputs and submits the issue in Japanese, "Customer feedback is delayed. I am very upset.", the input and emotions are recognized as "confused" and "stressed." The server translates the received input into English, resulting in "Customer feedback is delayed. I am very upset." The grammar checker corrects the input to make it grammatically correct, and then stores it in the database. Next, it is aggregated with similar input from other employees, and the generative AI creates a summary: "Several employees report delays in customer feedback and are very stressed." This summary and emotional information are sent to the administrator interface, where administrators can make quick decisions based on it.
[0747] In this way, the entire system functions efficiently, supporting information sharing among employees who speak different languages and quick decision-making that takes emotional information into account.
[0748] Example 2
[0749] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0750] There is a need for a system that can efficiently collect, translate, check grammar, and analyze emotions from users who speak different languages, and then report the results to administrators. This system must also be able to accurately analyze users' emotional information and generate summaries that take this information into account, thereby solving the problem of providing administrators with information that includes the users' emotional state.
[0751] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0752] In this invention, the server includes means for receiving information input in different languages, means for translating the received information into a common language, means for checking and correcting the grammar of the translated information, means for centrally managing multiple pieces of information, means for summarizing the centrally managed information, means for reporting the summarized information to an administrator, means for collecting emotional information from the input information, means for analyzing the emotional information, and means for generating a summary that includes the emotional information. This makes it possible to efficiently collect and manage information from users who speak different languages and to provide information that includes their emotional states.
[0753] "Different languages" refers to languages used by users in multiple different countries or regions.
[0754] "Input information" refers to text data and assignments entered by the user using the terminal.
[0755] "Means of receiving" refers to the mechanism or method by which the server receives data sent from the terminal.
[0756] "Common language" refers to a standard language that is used uniformly within the system, such as English.
[0757] "Means of translation" refers to the functions and methods for converting information entered in different languages into a common language.
[0758] "Grammar checking and correction means" refers to a method or module for detecting grammatical errors in the translated text and correcting it to make it grammatically correct.
[0759] "Means of centralized management" refers to a method of integrating multiple input data and translation results and centrally managing them in a database or management system.
[0760] "Means of summarization" refers to techniques or methods, such as the use of generative AI models, to succinctly summarize large amounts of input information or data.
[0761] "Means for reporting to administrator" refers to the method or interface for notifying the administrator of the summarized information.
[0762] "Emotional information" refers to emotional states analyzed from information entered by the user, such as data on "confusion" or "stress."
[0763] "Means for collecting emotional information" refers to the technology or method for analyzing user input data to recognize and collect emotions.
[0764] "Means for analyzing emotional information" refers to technology for analyzing collected emotional data and evaluating its content and trends.
[0765] "Means for generating summaries that include emotional information" refers to techniques and methods for creating summaries of input data that take emotional information into account.
[0766] The present invention is a system that receives information entered in different languages, translates it into a common language, and checks and corrects the grammar of the information. This system also has the function of collecting and analyzing emotional information, centrally managing the collected information, and finally summarizing and reporting it to an administrator. Specific embodiments of this system are described below.
[0767] Users input business issues and information into the device in their native language. For example, they can enter something like "Feedback from customers is delayed. We are having a lot of trouble" in Japanese and press the "Send" button. Users can first select the language they want to use and then enter the issue or information in the text field.
[0768] The device receives the user's input data. At the same time, a built-in emotion engine simultaneously collects emotional information based on the input content and the user's input speed. This emotion engine analyzes and recognizes emotions such as "confusion" or "stress" from the user's text content.
[0769] The server receives the user's input data and emotional information sent from the device. The received data is first passed to a multilingual translation AI module, which translates the different input languages (such as Japanese) into a common language (such as English). An industry-standard AI translation engine is often used as the translation module.
[0770] The translated text is then sent to a grammar checker module, which can incorporate existing grammar checking software such as Grammarly, to check for grammatical errors and correct them if necessary.
[0771] Once the translation and grammar check are complete, the information is centrally managed by a server. This information is stored in a database that aggregates input data from different employees. This database management system may use MySQL or PostgreSQL.
[0772] Next, the server aggregates multiple issues stored in the database and passes them to a generation AI to create a summary. Here, emotional information is also provided to the generation AI, and the tendency and intensity of emotions are reflected in the summarization process. For example, a summary is generated that includes emotional information, such as "Most employees are stressed by delays in feedback from customers."
[0773] Finally, the server sends the generated summary and emotion information to the administrator interface, which is updated in real time, allowing the administrator to instantly check the information. This allows the administrator to understand the employee's emotional state and provides information for better decision-making.
[0774] For example, if a user inputs an issue in Japanese such as "Customer feedback is delayed. I am very upset." and the input contains emotional expressions, the input is sent to the server via the terminal. The server translates this into English as "Customer feedback is delayed. I am very upset." A grammar checker then corrects the input to make it grammatically correct. The server also uses an emotion engine to recognize emotions such as "confused" and "stressed." Next, multiple similar reports are aggregated, and a generative AI creates a summary such as "Several employees report delays in customer feedback and are very stressed." Finally, this summary and emotional information are sent to the administrator interface, where the administrator reviews it and considers countermeasures.
[0775] An example prompt is, "Recognize the sentiment from the user's input data and summarize the text in a way that includes sentiment. Summarize the text as follows: 'Several employees report delays in customer feedback and are very stressed.'"
[0776] As described above, the present invention can provide a system that can efficiently manage information regardless of language differences and can report information including emotional information to a manager.
[0777] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0778] Step 1:
[0779] The user inputs the issue or information into the device in their native language. After inputting, they select the language they want to use, enter something like "Feedback from customers is delayed. We are having a lot of trouble," in the text field, and press the "Send" button. This saves the input information as text data on the device.
[0780] Step 2:
[0781] The device receives text data entered by the user. At the same time, the emotion engine analyzes the text content and behavioral data such as the user's typing speed, and collects and recognizes emotional information such as "confusion" or "stress." The input is text data, and the output is text data and emotional information. Specifically, the device also collects the user's keystroke data and sends it to the emotion engine for analysis.
[0782] Step 3:
[0783] The device sends the user's text data and emotional information to the server. Specifically, the device requests the server to send data using the HTTP protocol. The input is text data and emotional information, and the output is sent to the server.
[0784] Step 4:
[0785] The server receives text data and emotion information sent from the device. The received data is first passed to a multilingual translation AI module, which translates the input language (e.g., Japanese) into a common language (e.g., English). The input is Japanese text data, and the output is English text data. The server calls the translation engine to convert the data.
[0786] Step 5:
[0787] The server sends the translated text data to the grammar checker module, which checks and corrects grammatical errors. The input is English text data, and the output is English text data with corrected grammar. The server passes the data to the grammar checker module and receives the corrected results.
[0788] Step 6:
[0789] The server stores the corrected text data and emotion information in a database for centralized management. The input is the corrected text data and emotion information, and the output is saving to the database. The server inserts data into the database management system using SQL queries.
[0790] Step 7:
[0791] The server aggregates multiple issues and emotion information stored in the database and passes it to the generative AI model to create a summary. The input is multiple data in the database, and the output is summary data. Specifically, the server retrieves the data using an SQL query and provides it to the generative AI model along with a prompt statement to generate a summary. An example of a prompt statement is "Several employees report delays in customer feedback and are very stressed."
[0792] Step 8:
[0793] The server sends the generated summary and emotion information to the administrator interface. The input is summary data and emotion information, and the output is interface updates. The server sends data to the administrator's browser in real time using WebSocket or HTTP protocol and updates the interface.
[0794] Through these steps, the system can efficiently translate, check grammar, analyze sentiment, and aggregate information in different languages to generate summaries and provide them to administrators.
[0795] (Application example 2)
[0796] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0797] In the past, when communicating between different languages, language translation, grammar checking, and emotion recognition tended to be performed in separate systems, making it difficult to centrally manage and summarize information and report it to managers immediately.In addition, in physical stores, there was a lack of means to receive feedback from multilingual customers in real time and process it smoothly, which placed a heavy burden on staff.
[0798] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving information input in different languages, means for translating the received information into a common language, means for checking and correcting the grammar of the translated information, means for recognizing the user's emotions, means for centrally managing multiple pieces of information, means for summarizing the centrally managed information, means for reporting the summarized information to a manager, and components that allow store staff to input customer feedback on-site in real time and instantly check the translation results and emotional information. This centralizes feedback processing in a multilingual environment, making it possible to simultaneously manage customer feedback and the emotional state of staff.
[0799] The "means for receiving information input in a different language" is a function for receiving information input by a user in a different language at a terminal.
[0800] The "means for translating received information into a common language" is a function for automatically translating received information in multiple languages into a common language.
[0801] The "means for checking and correcting the grammar of translated information" is a function for detecting grammatical errors in translated information and automatically correcting them.
[0802] The "means for recognizing the user's emotions" is a function that analyzes the user's input content and behavior at the time of input and identifies the user's emotional state.
[0803] "Means for centrally managing multiple pieces of information" refers to a function that consolidates and manages multiple pieces of input information and related emotional information in one place.
[0804] "Means for summarizing centrally managed information" is a function that generates a short summary of the main points based on the aggregated information.
[0805] The "means for reporting summarized information to the administrator" is a function for instantly notifying the administrator of summarized information and emotional information.
[0806] "A component that allows store staff to input customer feedback in real time on-site and instantly check the translation results and emotional information" is a device that allows store staff to input feedback from customers in real time and instantly view it along with the translation and emotional analysis results.
[0807] This invention is a unified management system that receives customer feedback in real time in a multilingual environment in a brick-and-mortar store, translates it, checks grammar, and recognizes emotions. Specific embodiments of this system are described below.
[0808] System configuration
[0809] The system mainly consists of three main components: a server, a terminal, and a user.
[0810] Server: Translates data, recognizes emotions, checks grammar, and centralizes and summarizes information.
[0811] Terminal: A device where store clerks can input customer feedback in real time and check the translated results and sentiment analysis results. This can be a smartphone, smart glasses, or a head-mounted display.
[0812] Users: Includes customers and store clerks, where customers provide feedback and store clerks input it into the terminal.
[0813] Program Overview
[0814] The server receives information (feedback) entered in different languages and translates it into a common language. Translation services such as the Google Cloud Translation API are used for translation. The translated text is then checked for grammar using services such as the Grammarly API and corrected as necessary. Emotion recognition is performed using services such as the Microsoft Azure Emotion API to analyze the user's emotional state.
[0815] The received information and analyzed sentiment information are centrally managed on a server. Multiple feedback information is stored in a database and summarized using a generative AI model, allowing managers to efficiently grasp the key points.
[0816] Hardware and software used
[0817] Hardware:
[0818] Smartphone
[0819] Smart Glasses
[0820] head-mounted display
[0821] software:
[0822] Google Cloud Translation API (translation)
[0823] Microsoft Azure Emotion API (emotion recognition)
[0824] Grammarly API (Grammar Check)
[0825] Processing Flow
[0826] 1. User (store clerk): Enters customer feedback into a device such as a smartphone.
[0827] 2. Terminal: Sends the entered information to the server.
[0828] 3. Server:
[0829] Translate information.
[0830] Check and correct the grammar of the translated text.
[0831] Recognize user emotions.
[0832] 4. Server: Centralizes the translated and analyzed information, generates summaries, and reports to the administrator.
[0833] 5. Terminal: Displays translation results and sentiment analysis results in real time.
[0834] Specific examples
[0835] A customer provides feedback in Japanese, saying, "Customer feedback is delayed. I am very upset." The store clerk enters this into a terminal. This input is sent to the server, where it is translated into English, becoming, "Customer feedback is delayed. I am very upset." Grammar checks are then performed to ensure the correct form is achieved. Emotion recognition identifies emotions such as "confusion" and "stress," and the aggregated data generates a summary: "Several employees report delays in customer feedback and are very stressed." Finally, this summary and emotional information are reported to management.
[0836] Prompt Sentence Examples
[0837] "Please translate the feedback below, check for grammar, and recognize emotional information.
[0838] Language: Japanese
[0839] Feedback: "Customers are very unhappy"
[0840] As described above, it is possible to simultaneously manage customer feedback and the emotional state of staff in a multilingual environment, and a system is provided that can instantly translate and respond to communication between store staff and customers.
[0841] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0842] Step 1:
[0843] User (store clerk): The store clerk uses a device such as a smartphone or smart glasses to input customer feedback in real time. The input data is in text format, for example, "Customer feedback is delayed. We are very troubled." The input data is temporarily stored in the device.
[0844] Step 2:
[0845] Terminal: The terminal sends the input text data to the server, along with input language information (e.g., feedback in Japanese and its timestamp). A secure communication protocol (e.g., HTTPS) is used for this data transmission.
[0846] Step 3:
[0847] Server: The server receives the text data and first translates it into a common language (e.g., English) using the Google Cloud Translation API. The input is the original Japanese feedback, and the output is the translated data, such as "Customer feedback is delayed. I am very upset."
[0848] Step 4:
[0849] Server: After getting the translated data, we use the Grammarly API to perform a grammar check. Here, the input is the translated English text and the output is the grammatically corrected text. The corrections could be, for example, "Customer feedback has been delayed. I am very upset."
[0850] Step 5:
[0851] Server: After the grammar check is complete, emotion recognition is performed using the Microsoft Azure Emotion API. The input is the user's text data ("Customer feedback has been delayed. I am very upset."), and the output is the recognized emotion information (e.g., "confused" or "stressed").
[0852] Step 6:
[0853] Server: The data that has been translated, checked for grammar, and recognized for emotion is managed in a centralized database. The server stores this data and also assigns metadata such as timestamps to each piece of data. The input is the original feedback and various analyzed data, and the output is a centralized database entry.
[0854] Step 7:
[0855] Server: Multiple pieces of feedback stored in a database are summarized using a generative AI model. The input is a set of accumulated database entries, and the output is summary information, such as "Several employees report delays in customer feedback and are very stressed."
[0856] Step 8:
[0857] Server: Finally, the generated summary information and related sentiment information are sent to the administrator interface, allowing the administrator to check information based on staff status and customer feedback in real time. The input is the summarized data, and the output is the display results on the interface accessed by the administrator.
[0858] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0859] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0860] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0861] [Fourth embodiment]
[0862] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0863] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0864] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0865] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0866] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0867] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0868] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0869] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0870] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0871] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0872] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0873] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0874] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0875] As a concrete example of the present invention, we will explain a multilingual collaborative task aggregation AI service. This system consists of three main components: a server, a terminal, and a user, each of which has a specific function.
[0876] Handling User Input
[0877] Users input their business issues into the terminal in their native language. This input is done through a text field, and after the user has written the issue, they press the "Submit" button. This operation causes the terminal to send the user's input as data to the server.
[0878] Translation and Grammar Check
[0879] The server receives user input sent from the terminal. The received data is first passed to a multilingual translation AI module, which translates the input language into a common language (e.g., English). The translated text is then sent to a grammar checker module, where it is checked for grammatical errors and corrected if necessary.
[0880] Aggregation and Summarization
[0881] The server centralizes translated assignments collected from multiple users. This process involves aggregating each user's input into a database. The aggregated data is then passed to a generative AI module, which further extracts and summarizes key points. This generative AI module has learned from past data and patterns to efficiently and accurately summarize information.
[0882] Reporting to management
[0883] The server sends the generated summary to an administrator interface, which updates in real time, allowing the administrator to instantly view the summarized information and make quick decisions based on the summary.
[0884] Specific examples
[0885] If a user inputs the issue "Customer feedback is delayed" in Japanese, the input is sent to the server via the terminal. The server translates this input into English as "Customer feedback is delayed." The grammar checker then corrects it to make it grammatically correct. Next, multiple similar reports are aggregated and a generative AI creates a summary: "Several employees report delays in customer feedback." Finally, this summary is sent to the administrator interface, where it is reviewed and evaluated by the administrator.
[0886] In this way, the present invention helps multilingual employees efficiently gather information, translate, check grammar, aggregate data, and summarize it, and quickly report it to managers, thereby supporting smooth communication and quick decision-making across language barriers.
[0887] The processing flow will be explained below.
[0888] Step 1:
[0889] The user enters business issues and information into the terminal in their native language. After selecting the language they want to use and entering the issue in the text field, they press the "Submit" button.
[0890] Step 2:
[0891] The terminal sends the user's input to the server. The terminal sends the input data to the server as an HTTP request. This data includes the language information selected by the user and the assignment text.
[0892] Step 3:
[0893] The server passes the received information to a multilingual translation AI, which translates it into a common language (e.g., English). The server analyzes the received data, sends the text and language information to the translation API, and obtains the translation results.
[0894] Step 4:
[0895] The server passes the translated text to the grammar checker module to check and correct grammatical errors. The server sends the translated text to the grammar checker API to make it grammatically correct.
[0896] Step 5:
[0897] The server centralizes the translated assignments from all employees. The server stores the translated and grammar-checked text in a database and aggregates input from different employees.
[0898] Step 6:
[0899] The server passes the aggregated data to the generative AI to create a summary. The server extracts multiple issues stored in the database and passes them to the generative AI module to create a summary.
[0900] Step 7:
[0901] The server sends the generated summary to the administrator interface, where it displays the summary created by the generating AI in real time, allowing for immediate confirmation.
[0902] Step 8:
[0903] The administrator checks the summary on the management screen and makes a decision. The administrator evaluates the summary content through the interface and takes necessary action or makes a decision.
[0904] Through specific actions at each step, the entire system functions efficiently, supporting information sharing and quick decision-making among employees who speak different languages.
[0905] Example 1
[0906] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0907] In today's global companies, when employees who speak different languages work together, language barriers become an obstacle, making it difficult to communicate efficiently and make quick decisions. To address this issue, systems are needed that can efficiently translate information, check grammar, aggregate data, and summarize it. In existing systems, these processes are often performed manually, requiring time and effort. Furthermore, a lack of centralized information management makes it difficult for managers to quickly obtain the information they need.
[0908] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0909] In this invention, the server includes means for receiving information entered in different languages, means for translating the received information into a common language, means for checking and correcting the grammar of the translated information, means for centrally managing multiple pieces of information, means for summarizing the centrally managed information using a generative model, and means for reporting the summarized information to an administrator in real time. This enables smooth communication across language barriers and efficient aggregation and summarization of information. As a result, administrators can quickly obtain the information they need and make quick decisions.
[0910] "Means for receiving" refers to the function for importing information entered in different languages from a terminal into a server.
[0911] "Means for translating into a common language" refers to a function for converting received information from a different language into one pre-specified language (primarily English).
[0912] "Grammar checking and correction means" refers to functionality for detecting and, if necessary, correcting grammatical errors in the translated information.
[0913] "Means for centralized management" refers to a function for aggregating information received from multiple users into a central database and managing it in a unified manner.
[0914] "Means for summarizing using a generative model" refers to a function that uses a generative AI model based on aggregated information to extract important points and create a summary.
[0915] "Means of reporting in real time" refers to a function that allows the administrator to be notified of the generated summary immediately without any time lag.
[0916] MODE FOR CARRYING OUT THE INVENTION
[0917] This invention is a system for a multilingual collaborative task aggregation AI service. This system consists of three elements: a server, a terminal, and a user. Each element has a specific function and works together to solve business tasks.
[0918] Handling User Input
[0919] Users input business issues into the terminal in their native language. This input is done through a text field, and after the user describes the issue, they press the "Submit" button. This operation causes the terminal to generate data and send it to the server as JSON-formatted data. For example, if a user inputs "Customer feedback is delayed," this text is sent from the terminal to the server.
[0920] Translation Processing
[0921] The server receives the data sent from the device. It analyzes the received data and passes it to a multilingual translation AI module. This translation AI module may use, for example, the Google Translate API or the DeepL API. The input Japanese text is translated into English, the common language. For example, "Customer feedback is delayed" is translated to "Customer feedback is delayed."
[0922] Grammar Check
[0923] The translated text is sent to a grammar checker module, which can use the Grammarly API or LanguageTool API, for example. The grammar checker detects grammatical errors and corrects them if necessary. For example, "Customer feedback is delayed." is checked for grammar and any necessary corrections are made.
[0924] Data Aggregation
[0925] The server aggregates all the translated texts sent by users into a database. This process uses MySQL or PostgreSQL as a database management system (DBMS), which allows for centralized management of each user's input.
[0926] Data Summary
[0927] The aggregated data is passed to a generative AI module, which uses generative AI models such as OpenAI's GPT-3 and ChatGPT. The generative AI learns from past data and patterns, extracts key points, and summarizes the text. For example, if multiple users submit similar reports, the generative AI will generate a summary such as "Several employees report delays in customer feedback."
[0928] Reporting to management
[0929] The generated summary is sent to an administrator interface, where a dashboard is displayed that updates in real time, allowing administrators to instantly view the summarized information. For example, when an administrator opens the dashboard, they might see a summary that reads, "Several employees report delays in customer feedback."
[0930] Example prompt sentences
[0931] Below are some specific examples of prompt sentences to input to the generative AI model.
[0932] 1. The user types "Customer feedback is delayed" in Japanese into the device and presses the send button.
[0933] 2. The device sends this input to the server as JSON format data.
[0934] 3. The server passes the received data to the Google Translate API and translates it into English: "Customer feedback is delayed."
[0935] 4. Check the grammar of the translated text using the Grammarly API and make corrections if necessary.
[0936] 5. The server stores and aggregates all user data in a MySQL database.
[0937] 6. Pass the aggregated data to ChatGPT and generate a summary: “Several employees report delays in customer feedback.”
[0938] 7. The server sends a summary to the administrator's dashboard, where the administrator can instantly view it.
[0939] In this way, the present invention enables smooth communication in a multilingual environment, and supports efficient information gathering and summarization, as well as rapid decision-making.
[0940] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0941] Step 1:
[0942] The user inputs the issue into the terminal. The user writes the business issue in their native language into the text field on the terminal and presses the "Submit" button. This operation converts the entered text into JSON format by the terminal and sends it to the server. The input is the text "Customer feedback is delayed," and the output is JSON format data.
[0943] Step 2:
[0944] The server receives the data sent from the device. It analyzes the received JSON data and passes it to a multilingual translation AI module. This translation AI module uses, for example, a multilingual translation API. It receives JSON data as input and obtains translated text data as output. For example, the Japanese data "Customer feedback is delayed" is translated into the English data "Customer feedback is delayed."
[0945] Step 3:
[0946] The server sends the translated text to a grammar checker module. The grammar checker module uses, for example, a grammar check API. It receives the translated text data as input, detects grammatical errors, and outputs corrected text data as necessary. For example, the text "Customer feedback is delayed." is grammatically checked to ensure there are no errors.
[0947] Step 4:
[0948] The server aggregates all translated text submitted by users into a database. The database management system (DBMS) can be, for example, a relational database. It receives translated text data as input and outputs it as a centralized database record. For example, text submitted by multiple users can be aggregated into a single database.
[0949] Step 5:
[0950] The server passes the aggregated data to a generative AI module. This module uses a generative AI model that takes the aggregated data from the database as input, extracts key points, and outputs summarized text data. For example, multiple matching reports might be summarized as "Several employees report delays in customer feedback."
[0951] Step 6:
[0952] The server sends the generated summary to an administrator interface, which displays a dashboard that updates in real time, allowing administrators to instantly view the summarized information. The interface receives summarized text data as input and displays it as an administrator-accessible dashboard as output. For example, a real-time summary such as "Several employees report delays in customer feedback" appears on the administrator's dashboard.
[0953] (Application example 1)
[0954] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0955] In an environment where employees speak a wide variety of languages, it is necessary to quickly and efficiently collect, translate, and summarize business issues and provide accurate information to managers. Real-time problem reporting and rapid decision-making are particularly important in logistics centers. However, the overlap of different languages and manual information processing can lead to delays and errors in the collection, translation, and transmission of information.
[0956] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0957] In this invention, the server includes a means for receiving information input in different languages, a means for translating the received information into a common language, and a means for checking and correcting the grammar of the translated information. This provides a means for centrally managing multiple pieces of information, a means for summarizing the centrally managed information, and a means for reporting the summarized information to an administrator. This allows for the input of information in real time using a smart device, and for the input information to be translated, checked for grammar, and summarized, and then reported promptly to an administrator.
[0958] "Information entered in different languages" refers to text data entered by users in their various native languages.
[0959] A "common language" is a unified language (e.g., English) that is translated from different languages.
[0960] "Grammar checking and correction" is the process of making the translated text grammatically correct.
[0961] "Centralized management" refers to the process of aggregating multiple pieces of information into a central database and managing them in an integrated manner.
[0962] "Summarizing" refers to extracting important points from aggregated information and summarizing them concisely.
[0963] "Reporting to administrator" is the process of sending the processed information to an administrator interface so that the administrator can review it.
[0964] "Inputting information in real time using a smart device" refers to inputting information instantly from a mobile device such as a smartphone or tablet.
[0965] A "generative model" is an algorithm that uses machine learning techniques to summarize and generate information.
[0966] This paper describes a system that realizes a multilingual collaborative task aggregation AI service in a logistics center. This system consists of three main elements: a server, a terminal, and a user, each of which has a specific function.
[0967] Handling User Input
[0968] Users input their work-related tasks in their native language using a smartphone. This input is done through a text field, and after the user has written the task, they press the "Submit" button to complete the task. The device then sends the user's input as data to the server.
[0969] Translation and Grammar Check
[0970] The server receives user input sent from the device. The received data is first translated into a common language (e.g., English) using the Google Translator API. This translated text is then sent to the grammar checking module via OpenAI's API, where grammatical errors are checked and corrected.
[0971] Aggregation and Summarization
[0972] The server centrally manages translated assignments collected from multiple users and aggregates them into a database. The aggregated data is then extracted and summarized by OpenAI's generative AI model. This generative AI model learns from past data and patterns to efficiently and accurately summarize information.
[0973] Report to the administrator
[0974] The server sends the generated summary to an administrator interface, which updates in real time, allowing administrators to instantly see the summarized information and make quick decisions.
[0975] Specific examples
[0976] If a user inputs an issue in Japanese such as "Errors in the inventory management system occur frequently," the input is sent to the server via the terminal. The server translates this input into English using the Google Translator API, resulting in "Errors in the inventory management system occur frequently." The translated text is then checked for grammar using the OpenAI API, and corrected to "Errors in the inventory management system are occurring frequently." The generation AI generates a summary from many similar reports: "Frequent errors in inventory system." This summary is sent to the administrator interface, where the administrator can take prompt action based on it.
[0977] Example prompt sentence:
[0978] Grammar check prompt: "Please correct the following text: Errors in the inventory management system occur frequently."
[0979] Summary prompt: "Summarize the following text: Errors in the inventory management system are occurring frequently."
[0980] In this way, the present invention enables multilingual employees to efficiently gather information, translate, check grammar, aggregate data, and summarize, and quickly report to management, thereby supporting smooth communication and quick decision-making across language barriers in logistics centers.
[0981] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0982] Step 1:
[0983] A user uses a smart device to input a business issue in their native language. Specifically, the user writes the issue in a text field on the smartphone application and presses the "Submit" button. The input issue (e.g., "Errors are occurring frequently in the inventory management system") is sent from the device to the server. The input data is text data in the user's native language.
[0984] Step 2:
[0985] The server receives user input sent from the device. The received data is first sent to the Google Translator API, where the native text data is translated into a common language (e.g., English). The translated text data (e.g., "Errors in the inventory management system occur frequently.") is generated.
[0986] Step 3:
[0987] The server sends the translated text data to OpenAI's API, where a grammar check is performed using a generative AI model. Specifically, the server detects and corrects grammatical errors in the translated text, generating corrected text data (e.g., "Errors in the inventory management system are occurring frequently.").
[0988] Step 4:
[0989] The server stores the grammar-checked text data in a database and manages it centrally. Multiple similar data sets are aggregated, and the text data is stored in a management database. This ensures that assignments from multiple users are managed in an integrated manner.
[0990] Step 5:
[0991] The server sends the aggregated data to OpenAI's generative AI model, which extracts key points and creates a summary. Specifically, the server selects key content from multiple translated and corrected text data and generates a summary text (e.g., "Frequent errors in inventory system.").
[0992] Step 6:
[0993] The server sends the generated summary text to an administrator interface, where administrators can view the summarized information using an interface that updates in real time, providing administrators with accurate information for quick decision making.
[0994] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0995] The present invention is a system that incorporates an emotion engine that recognizes the user's emotions, as well as a means for receiving information entered in different languages, translating it into a common language, and checking and correcting the grammar of the information. This system consists of three main components: a server, a terminal, and a user, each of which has a specific function.
[0996] User Input Processing and Emotion Recognition
[0997] Users input work-related issues and information into the device in their native language. After selecting the language they want to use and entering the issue in the text field, they press the "Submit" button. The device then sends this input data to the server. During this process, the device also collects emotional information along with the user's input data. The emotional information is recognized by an emotion engine that analyzes the user's input content and behavior during input.
[0998] Translation and Grammar Check
[0999] The server receives user input and emotion information sent from the device. The received data is first passed to a multilingual translation AI module, which translates the input language into a common language (e.g., English). The translated text is then sent to a grammar checker module, where it is checked for grammatical errors and corrected if necessary.
[1000] Unified management of emotional information
[1001] The server centrally manages the emotional information along with the translated and grammar-checked information. The server stores this information in a database and aggregates input data from different employees. At this time, the emotional information is treated like regular text data and stored in the database.
[1002] Aggregation and Summarization
[1003] The server aggregates multiple issues stored in the database and passes them to the generation AI to create a summary. Emotional information is also provided to the generation AI, and the tendency and intensity of emotions are reflected in the summarization process. For example, if many employees are feeling stressed, that emotional information will be noted as part of the summary.
[1004] Report to the administrator
[1005] The server sends the generated summary and emotion information to the administrator interface, which is updated in real time, allowing the administrator to instantly check the summary and emotion information. This allows the administrator to understand not only the text information but also the emotional state of employees, enabling them to make better decisions.
[1006] Specific examples
[1007] If a user inputs an issue in Japanese, such as "Customer feedback is delayed. I am very upset," and the input includes emotional expressions, the input is sent to the server via the terminal. The server translates this into English as "Customer feedback is delayed. I am very upset." A grammar checker then corrects the input to make it grammatically correct. The server also uses an emotion engine to recognize emotions such as "confused" and "stressed." Next, multiple similar reports are aggregated, and a generative AI creates a summary: "Several employees report delays in customer feedback and are very stressed." Finally, this summary and emotional information are sent to the administrator interface, where the administrator reviews it and considers countermeasures.
[1008] This system integrates multilingual information collection, translation, grammar checking, and emotion recognition, and can centrally manage and summarize data for quick reporting, eliminating communication barriers caused by language and cultural differences and supporting appropriate decision-making that takes into account the emotional state of employees.
[1009] The processing flow will be explained below.
[1010] Step 1:
[1011] The user enters work-related issues and information into the device in their native language. After selecting the language they want to use and entering the issue in the text field, they press the "Submit" button. This operation also collects emotional information.
[1012] Step 2:
[1013] The device sends the user's input and emotional information to the server. The input content and emotional information are sent to the server as an HTTP request. The emotional information is the result of an emotion engine analyzing the user's input content and behavior at the time of input.
[1014] Step 3:
[1015] The server passes the received information to a multilingual translation AI, which translates it into a common language (e.g., English). The server analyzes the received data, sends the text and language information to the translation API, and obtains the translation results.
[1016] Step 4:
[1017] The server passes the translated text to the grammar checker module, which checks and corrects grammatical errors, and the translation result is sent to the grammar checker API, which formats it to be grammatically correct.
[1018] Step 5:
[1019] The server centralizes the translated and grammar-checked information, as well as the sentiment information, which is then stored in a database that aggregates input from different employees.
[1020] Step 6:
[1021] The server passes the collected data to the generation AI to create a summary. Using the issue information and emotion information stored in the database, the generation AI module extracts important points and reflects them in the summary.
[1022] Step 7:
[1023] The server sends the generated summary and emotion information to the administrator interface, which is updated in real time, allowing the administrator to instantly check the summary and emotion information.
[1024] Step 8:
[1025] The administrator checks the summary on the management screen and makes a decision. The administrator evaluates the summary content and sentiment information through the interface and quickly takes necessary actions and makes decisions.
[1026] As a specific example, if a user inputs and submits the issue in Japanese, "Customer feedback is delayed. I am very upset.", the input and emotions are recognized as "confused" and "stressed." The server translates the received input into English, resulting in "Customer feedback is delayed. I am very upset." The grammar checker corrects the input to make it grammatically correct, and then stores it in the database. Next, it is aggregated with similar input from other employees, and the generative AI creates a summary: "Several employees report delays in customer feedback and are very stressed." This summary and emotional information are sent to the administrator interface, where administrators can make quick decisions based on it.
[1027] In this way, the entire system functions efficiently, supporting information sharing among employees who speak different languages and quick decision-making that takes emotional information into account.
[1028] Example 2
[1029] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1030] There is a need for a system that can efficiently collect, translate, check grammar, and analyze emotions from users who speak different languages, and then report the results to administrators. This system must also be able to accurately analyze users' emotional information and generate summaries that take this information into account, thereby solving the problem of providing administrators with information that includes the users' emotional state.
[1031] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1032] In this invention, the server includes means for receiving information input in different languages, means for translating the received information into a common language, means for checking and correcting the grammar of the translated information, means for centrally managing multiple pieces of information, means for summarizing the centrally managed information, means for reporting the summarized information to an administrator, means for collecting emotional information from the input information, means for analyzing the emotional information, and means for generating a summary that includes the emotional information. This makes it possible to efficiently collect and manage information from users who speak different languages and to provide information that includes their emotional states.
[1033] "Different languages" refers to languages used by users in multiple different countries or regions.
[1034] "Input information" refers to text data and assignments entered by the user using the terminal.
[1035] "Means of receiving" refers to the mechanism or method by which the server receives data sent from the terminal.
[1036] "Common language" refers to a standard language that is used uniformly within the system, such as English.
[1037] "Means of translation" refers to the functions and methods for converting information entered in different languages into a common language.
[1038] "Grammar checking and correction means" refers to a method or module for detecting grammatical errors in the translated text and correcting it to make it grammatically correct.
[1039] "Means of centralized management" refers to a method of integrating multiple input data and translation results and centrally managing them in a database or management system.
[1040] "Means of summarization" refers to techniques or methods, such as the use of generative AI models, to succinctly summarize large amounts of input information or data.
[1041] "Means for reporting to administrator" refers to the method or interface for notifying the administrator of the summarized information.
[1042] "Emotional information" refers to emotional states analyzed from information entered by the user, such as data on "confusion" or "stress."
[1043] "Means for collecting emotional information" refers to the technology or method for analyzing user input data to recognize and collect emotions.
[1044] "Means for analyzing emotional information" refers to technology for analyzing collected emotional data and evaluating its content and trends.
[1045] "Means for generating summaries that include emotional information" refers to techniques and methods for creating summaries of input data that take emotional information into account.
[1046] The present invention is a system that receives information entered in different languages, translates it into a common language, and checks and corrects the grammar of the information. This system also has the function of collecting and analyzing emotional information, centrally managing the collected information, and finally summarizing and reporting it to an administrator. Specific embodiments of this system are described below.
[1047] Users input business issues and information into the device in their native language. For example, they can enter something like "Feedback from customers is delayed. We are having a lot of trouble" in Japanese and press the "Send" button. Users can first select the language they want to use and then enter the issue or information in the text field.
[1048] The device receives the user's input data. At the same time, a built-in emotion engine simultaneously collects emotional information based on the input content and the user's input speed. This emotion engine analyzes and recognizes emotions such as "confusion" or "stress" from the user's text content.
[1049] The server receives the user's input data and emotional information sent from the device. The received data is first passed to a multilingual translation AI module, which translates the different input languages (such as Japanese) into a common language (such as English). An industry-standard AI translation engine is often used as the translation module.
[1050] The translated text is then sent to a grammar checker module, which can incorporate existing grammar checking software such as Grammarly, to check for grammatical errors and correct them if necessary.
[1051] Once the translation and grammar check are complete, the information is centrally managed by a server. This information is stored in a database that aggregates input data from different employees. This database management system may use MySQL or PostgreSQL.
[1052] Next, the server aggregates multiple issues stored in the database and passes them to a generation AI to create a summary. Here, emotional information is also provided to the generation AI, and the tendency and intensity of emotions are reflected in the summarization process. For example, a summary is generated that includes emotional information, such as "Most employees are stressed by delays in feedback from customers."
[1053] Finally, the server sends the generated summary and emotion information to the administrator interface, which is updated in real time, allowing the administrator to instantly check the information. This allows the administrator to understand the employee's emotional state and provides information for better decision-making.
[1054] For example, if a user inputs an issue in Japanese such as "Customer feedback is delayed. I am very upset." and the input contains emotional expressions, the input is sent to the server via the terminal. The server translates this into English as "Customer feedback is delayed. I am very upset." A grammar checker then corrects the input to make it grammatically correct. The server also uses an emotion engine to recognize emotions such as "confused" and "stressed." Next, multiple similar reports are aggregated, and a generative AI creates a summary such as "Several employees report delays in customer feedback and are very stressed." Finally, this summary and emotional information are sent to the administrator interface, where the administrator reviews it and considers countermeasures.
[1055] An example prompt is, "Recognize the sentiment from the user's input data and summarize the text in a way that includes sentiment. Summarize the text as follows: 'Several employees report delays in customer feedback and are very stressed.'"
[1056] As described above, the present invention can provide a system that can efficiently manage information regardless of language differences and can report information including emotional information to a manager.
[1057] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1058] Step 1:
[1059] The user inputs the issue or information into the device in their native language. After inputting, they select the language they want to use, enter something like "Feedback from customers is delayed. We are having a lot of trouble," in the text field, and press the "Send" button. This saves the input information as text data on the device.
[1060] Step 2:
[1061] The device receives text data entered by the user. At the same time, the emotion engine analyzes the text content and behavioral data such as the user's typing speed, and collects and recognizes emotional information such as "confusion" or "stress." The input is text data, and the output is text data and emotional information. Specifically, the device also collects the user's keystroke data and sends it to the emotion engine for analysis.
[1062] Step 3:
[1063] The device sends the user's text data and emotional information to the server. Specifically, the device requests the server to send data using the HTTP protocol. The input is text data and emotional information, and the output is sent to the server.
[1064] Step 4:
[1065] The server receives text data and emotion information sent from the device. The received data is first passed to a multilingual translation AI module, which translates the input language (e.g., Japanese) into a common language (e.g., English). The input is Japanese text data, and the output is English text data. The server calls the translation engine to convert the data.
[1066] Step 5:
[1067] The server sends the translated text data to the grammar checker module, which checks and corrects grammatical errors. The input is English text data, and the output is English text data with corrected grammar. The server passes the data to the grammar checker module and receives the corrected results.
[1068] Step 6:
[1069] The server stores the corrected text data and emotion information in a database for centralized management. The input is the corrected text data and emotion information, and the output is saving to the database. The server inserts data into the database management system using SQL queries.
[1070] Step 7:
[1071] The server aggregates multiple issues and emotion information stored in the database and passes it to the generative AI model to create a summary. The input is multiple data in the database, and the output is summary data. Specifically, the server retrieves the data using an SQL query and provides it to the generative AI model along with a prompt statement to generate a summary. An example of a prompt statement is "Several employees report delays in customer feedback and are very stressed."
[1072] Step 8:
[1073] The server sends the generated summary and emotion information to the administrator interface. The input is summary data and emotion information, and the output is interface updates. The server sends data to the administrator's browser in real time using WebSocket or HTTP protocol and updates the interface.
[1074] Through these steps, the system can efficiently translate, check grammar, analyze sentiment, and aggregate information in different languages to generate summaries and provide them to administrators.
[1075] (Application example 2)
[1076] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1077] In the past, when communicating between different languages, language translation, grammar checking, and emotion recognition tended to be performed in separate systems, making it difficult to centrally manage and summarize information and report it to managers immediately.In addition, in physical stores, there was a lack of means to receive feedback from multilingual customers in real time and process it smoothly, which placed a heavy burden on staff.
[1078] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving information input in different languages, means for translating the received information into a common language, means for checking and correcting the grammar of the translated information, means for recognizing the user's emotions, means for centrally managing multiple pieces of information, means for summarizing the centrally managed information, means for reporting the summarized information to a manager, and components that allow store staff to input customer feedback on-site in real time and instantly check the translation results and emotional information. This centralizes feedback processing in a multilingual environment, making it possible to simultaneously manage customer feedback and the emotional state of staff.
[1079] The "means for receiving information input in a different language" is a function for receiving information input by a user in a different language at a terminal.
[1080] The "means for translating received information into a common language" is a function for automatically translating received information in multiple languages into a common language.
[1081] The "means for checking and correcting the grammar of translated information" is a function for detecting grammatical errors in translated information and automatically correcting them.
[1082] The "means for recognizing the user's emotions" is a function that analyzes the user's input content and behavior at the time of input and identifies the user's emotional state.
[1083] "Means for centrally managing multiple pieces of information" refers to a function that consolidates and manages multiple pieces of input information and related emotional information in one place.
[1084] "Means for summarizing centrally managed information" is a function that generates a short summary of the main points based on the aggregated information.
[1085] The "means for reporting summarized information to the administrator" is a function for instantly notifying the administrator of summarized information and emotional information.
[1086] "A component that allows store staff to input customer feedback in real time on-site and instantly check the translation results and emotional information" is a device that allows store staff to input feedback from customers in real time and instantly view it along with the translation and emotional analysis results.
[1087] This invention is a unified management system that receives customer feedback in real time in a multilingual environment in a brick-and-mortar store, translates it, checks grammar, and recognizes emotions. Specific embodiments of this system are described below.
[1088] System configuration
[1089] The system mainly consists of three main components: a server, a terminal, and a user.
[1090] Server: Translates data, recognizes emotions, checks grammar, and centralizes and summarizes information.
[1091] Terminal: A device where store clerks can input customer feedback in real time and check the translated results and sentiment analysis results. This can be a smartphone, smart glasses, or a head-mounted display.
[1092] Users: Includes customers and store clerks, where customers provide feedback and store clerks input it into the terminal.
[1093] Program Overview
[1094] The server receives information (feedback) entered in different languages and translates it into a common language. Translation services such as the Google Cloud Translation API are used for translation. The translated text is then checked for grammar using services such as the Grammarly API and corrected as necessary. Emotion recognition is performed using services such as the Microsoft Azure Emotion API to analyze the user's emotional state.
[1095] The received information and analyzed sentiment information are centrally managed on a server. Multiple feedback information is stored in a database and summarized using a generative AI model, allowing managers to efficiently grasp the key points.
[1096] Hardware and software used
[1097] Hardware:
[1098] Smartphone
[1099] Smart Glasses
[1100] head-mounted display
[1101] software:
[1102] Google Cloud Translation API (translation)
[1103] Microsoft Azure Emotion API (emotion recognition)
[1104] Grammarly API (Grammar Check)
[1105] Processing Flow
[1106] 1. User (store clerk): Enters customer feedback into a device such as a smartphone.
[1107] 2. Terminal: Sends the entered information to the server.
[1108] 3. Server:
[1109] Translate information.
[1110] Check and correct the grammar of the translated text.
[1111] Recognize user emotions.
[1112] 4. Server: Centralizes the translated and analyzed information, generates summaries, and reports to the administrator.
[1113] 5. Terminal: Displays translation results and sentiment analysis results in real time.
[1114] Specific examples
[1115] A customer provides feedback in Japanese, saying, "Customer feedback is delayed. I am very upset." The store clerk enters this into a terminal. This input is sent to the server, where it is translated into English, becoming, "Customer feedback is delayed. I am very upset." Grammar checks are then performed to ensure the correct form is achieved. Emotion recognition identifies emotions such as "confusion" and "stress," and the aggregated data generates a summary: "Several employees report delays in customer feedback and are very stressed." Finally, this summary and emotional information are reported to management.
[1116] Prompt Sentence Examples
[1117] "Please translate the feedback below, check for grammar, and recognize emotional information.
[1118] Language: Japanese
[1119] Feedback: "Customers are very unhappy"
[1120] As described above, it is possible to simultaneously manage customer feedback and the emotional state of staff in a multilingual environment, and a system is provided that can instantly translate and respond to communication between store staff and customers.
[1121] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1122] Step 1:
[1123] User (store clerk): The store clerk uses a device such as a smartphone or smart glasses to input customer feedback in real time. The input data is in text format, for example, "Customer feedback is delayed. We are very troubled." The input data is temporarily stored in the device.
[1124] Step 2:
[1125] Terminal: The terminal sends the input text data to the server, along with input language information (e.g., feedback in Japanese and its timestamp). A secure communication protocol (e.g., HTTPS) is used for this data transmission.
[1126] Step 3:
[1127] Server: The server receives the text data and first translates it into a common language (e.g., English) using the Google Cloud Translation API. The input is the original Japanese feedback, and the output is the translated data, such as "Customer feedback is delayed. I am very upset."
[1128] Step 4:
[1129] Server: After getting the translated data, we use the Grammarly API to perform a grammar check. Here, the input is the translated English text and the output is the grammatically corrected text. The corrections could be, for example, "Customer feedback has been delayed. I am very upset."
[1130] Step 5:
[1131] Server: After the grammar check is complete, emotion recognition is performed using the Microsoft Azure Emotion API. The input is the user's text data ("Customer feedback has been delayed. I am very upset."), and the output is the recognized emotion information (e.g., "confused" or "stressed").
[1132] Step 6:
[1133] Server: The data that has been translated, checked for grammar, and recognized for emotion is managed in a centralized database. The server stores this data and also assigns metadata such as timestamps to each piece of data. The input is the original feedback and various analyzed data, and the output is a centralized database entry.
[1134] Step 7:
[1135] Server: Multiple pieces of feedback stored in a database are summarized using a generative AI model. The input is a set of accumulated database entries, and the output is summary information, such as "Several employees report delays in customer feedback and are very stressed."
[1136] Step 8:
[1137] Server: Finally, the generated summary information and related sentiment information are sent to the administrator interface, allowing the administrator to check information based on staff status and customer feedback in real time. The input is the summarized data, and the output is the display results on the interface accessed by the administrator.
[1138] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1139] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1140] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1141] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1142] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1143] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1144] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1145] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1146] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1147] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1148] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1149] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1150] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1151] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1152] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1153] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1154] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1155] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1156] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1157] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1158] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1159] The following is further disclosed regarding the above embodiment.
[1160] (Claim 1)
[1161] means for receiving information entered in different languages;
[1162] means for translating the received information into a common language;
[1163] means for checking and correcting the grammar of the translated information;
[1164] A means for centrally managing multiple pieces of information;
[1165] A means of summarizing centralized information;
[1166] The system includes a means for reporting the summarized information to management.
[1167] (Claim 2)
[1168] 10. The system of claim 1, wherein the means for summarizing the translated information uses a generative model.
[1169] (Claim 3)
[1170] 10. The system of claim 1, which receives and translates information entered in different languages in real time.
[1171] "Example 1"
[1172] (Claim 1)
[1173] means for receiving information entered in different languages;
[1174] means for translating the received information into a common language;
[1175] a means of checking and correcting the grammar of the translated information;
[1176] A means for centrally managing multiple pieces of information;
[1177] A means for summarizing centrally managed information using a generative model;
[1178] A system that includes a means for reporting summarized information to management in real time.
[1179] (Claim 2)
[1180] 10. The system of claim 1, wherein the means for summarizing the translated information uses a generative model.
[1181] (Claim 3)
[1182] 10. The system of claim 1, which receives and translates information entered in different languages in real time.
[1183] "Application Example 1"
[1184] (Claim 1)
[1185] means for receiving information entered in different languages;
[1186] means for translating the received information into a common language;
[1187] means for checking and correcting the grammar of the translated information;
[1188] A means for centrally managing multiple pieces of information;
[1189] A means of summarizing centralized information;
[1190] a means for reporting the summarized information to management;
[1191] a means for inputting information in real time using a smart device;
[1192] A system that includes the means to translate, check grammar, summarize, and report to an administrator based on the input information.
[1193] (Claim 2)
[1194] 10. The system of claim 1, wherein the means for summarizing the translated information uses a generative model.
[1195] (Claim 3)
[1196] 10. The system of claim 1, which receives and translates information entered in different languages in real time.
[1197] "Example 2: Combining Emotion Engines"
[1198] (Claim 1)
[1199] means for receiving information entered in different languages;
[1200] means for translating the received information into a common language;
[1201] means for checking and correcting the grammar of the translated information;
[1202] A means for centrally managing multiple pieces of information;
[1203] A means of summarizing centralized information;
[1204] a means for reporting the summarized information to management;
[1205] A means for collecting emotion information from input information;
[1206] A means for analyzing emotional information;
[1207] A system including means for generating a summary in a manner that includes emotional information.
[1208] (Claim 2)
[1209] 10. The system of claim 1, wherein the summarized information and sentiment information are generated using a generative model.
[1210] (Claim 3)
[1211] 10. The system of claim 1, wherein the system receives and translates information and emotion information input in different languages in real time.
[1212] "Application example 2 when combining emotion engines"
[1213] (Claim 1)
[1214] means for receiving information entered in different languages;
[1215] means for translating the received information into a common language;
[1216] means for checking and correcting the grammar of the translated information;
[1217] means for recognizing a user's emotion;
[1218] A means for centrally managing multiple pieces of information;
[1219] A means of summarizing centralized information;
[1220] a means for reporting the summarized information to management;
[1221] A component that allows store staff to input customer feedback in real time on-site and instantly check the translation results and emotional information
[1222] A system including:
[1223] (Claim 2)
[1224] 10. The system of claim 1, wherein the means for summarizing the translated information uses a generative model.
[1225] (Claim 3)
[1226] 10. The system of claim 1, which receives and translates information and emotion information input in different languages in real time. [Explanation of symbols]
[1227] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving information entered in different languages; means for translating the received information into a common language; means for checking and correcting the grammar of the translated information; A means for centrally managing multiple pieces of information; A means of summarizing centralized information; The system includes a means for reporting the summarized information to management.
2. 10. The system of claim 1, wherein the means for summarizing the translated information uses a generative model.
3. 10. The system of claim 1, wherein the system receives and translates information entered in different languages in real time.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A