System

An integrated communication support system using a generative AI model addresses the inefficiencies of conventional tools by allowing users to perform translation, summarization, and creative idea generation within a single service, improving user experience and flexibility.

JP2026028906APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024131523
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Conventional communication tools struggle to efficiently meet diverse needs such as foreign language translation, summarization, and creative idea generation, requiring users to utilize multiple services separately, which hinders efficient communication.

Method used

An integrated communication support system utilizing a generative AI model that allows users to input text, receive and process it for translation, summarization, or creative idea generation, and convert the output into speech, all within a single service.

Benefits of technology

Enables users to efficiently meet various needs like translation, summarization, and creative idea suggestions through an integrated system, enhancing user experience and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026028906000001_ABST
    Figure 2026028906000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for a user to input text; means for receiving and forwarding the text to a generative AI model for analysis; means for the generative AI model to generate an appropriate response based on the text; means for converting the generated response to audio; and means for providing the audio to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional communication tools have had difficulty meeting diverse needs, such as foreign language translation, summarizing long texts, and generating new ideas. In particular, these processes are performed separately, forcing users to use multiple services, hindering efficient communication. The present invention aims to solve these problems by providing an integrated communication support system that utilizes generative AI models. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for a user to input text, a means for receiving the text and transmitting it to a generative AI model for analysis, a means for the generative AI model to generate an appropriate response based on the text, a means for converting the generated response into speech, and a means for providing the speech to the user. The system is capable of performing a variety of tasks, such as translation, summarization, and creative idea generation, and is designed to allow users to efficiently use these functions within a single service.

[0006] "User" refers to an individual or organization that uses the system.

[0007] "Text" refers to sentences or characters of characters that a user inputs into the system.

[0008] An "input means" is an interface or device through which a user provides text to a system.

[0009] "Means for receiving" refers to a method or device by which the system captures text entered by a user.

[0010] "Means for transmission to the generative AI model for analysis" means the process or method within the system for passing received text to the generative AI model for appropriate processing.

[0011] A "generative AI model" refers to artificial intelligence that uses machine learning and deep learning techniques to generate appropriate responses to user text.

[0012] "Means for generating appropriate responses" refers to how a generative AI model generates appropriate output, such as a translation, summary, or creative suggestion, based on the text it receives from the user.

[0013] "Means for converting to voice" refers to technology or devices for converting the generated text response into voice data.

[0014] "Means for providing" refers to a method or device for conveying data converted into voice to a user.

[0015] "System" refers to a series of processes or devices that are configured by combining the above means. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention relates to a communication support system that processes user-entered text using a generative AI model, translating, summarizing, and proposing creative ideas, and then providing the results as voice data. This system is comprised of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0038] The user inputs text through their device. The device has a text input field, and when the user inputs text, the text is received by the device. The received text is formatted into a standard data format such as JSON and then sent to the server.

[0039] The server analyzes the data received from the device and determines whether it corresponds to a translation task, a summarization task, or a creative idea suggestion. The generative AI model is then used for this processing. Specifically, the server sends an analysis request to the generative AI model, which then translates, summarizes, or generates creative ideas as needed.

[0040] The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. For example, when a user enters the text "Translate 'Hello' to Japanese," the generative AI model generates the translation result "Hello." This generated response text is then reformatted into JSON format and returned to the server.

[0041] The server receives the response text from the generative AI model and sends it to the device. The device uses speech synthesis technology to convert the received text into audio data. The text is converted into audio through a speech synthesis API, and the audio data is provided to the user through the speaker.

[0042] For example, if a user types "This is a long document that needs summarizing...", the generative AI model will summarize the long sentence and generate the summary text "This document needs summarizing". The device will receive this summary text and play it back to the user via speech synthesis.

[0043] If a user types "Give me an idea for a science project," the generative AI model will generate creative ideas, such as "Build a small wind turbine using household materials." The device will convert this idea into audio data and play it back to the user.

[0044] This allows users to efficiently meet a variety of needs, such as translation, summarization, and creative idea suggestions.The system of the present invention provides multiple functions in an integrated manner to support user communication, thereby resolving conventional problems.

[0045] The processing flow will be explained below.

[0046] Step 1:

[0047] A user enters text into an input field on a terminal. For example, "Translate 'Hello' to Japanese."

[0048] Step 2:

[0049] The device captures the user's input text and formats it into JSON format. For example, it generates data like {"action": "translate", "text": "Hello", "target_language": "Japanese"}.

[0050] Step 3:

[0051] The device sends the formatted data to the server, using an Internet connection to send a request to the server.

[0052] Step 4:

[0053] The server analyzes the data received from the device. It parses the received JSON data and checks that the "action" field is "translate".

[0054] Step 5:

[0055] The server transfers the analyzed data to the generative AI model. Specifically, it calls the generative AI model's API and requests a translation task. For example, make the following API call:

[0056] python

[0057] translation_result = ai_model.translate(text="Hello", target_language="Japanese")

[0058] Step 6:

[0059] The generative AI model translates the text "Hello" to "Hello." The translation result is returned to the server.

[0060] Step 7:

[0061] The server formats the translation results received from the AI ​​model into JSON format. For example, it generates data like {"translated_text": "Hello"}.

[0062] Step 8:

[0063] The server sends the formatted data to the device, again using the internet connection to return a response to the device.

[0064] Step 9:

[0065] The device parses the JSON data received from the server, interpreting the format {"translated_text": "Hello"} and extracting the contents of the "translated_text" field.

[0066] Step 10:

[0067] Convert text received by the device into speech. Use a speech synthesis API to convert text into speech data. Example: Use the following API:

[0068] python

[0069] synthesized_audio = text_to_speech("Hello")

[0070] Step 11:

[0071] The terminal plays the generated voice data to the user, and provides the voice to the user through the terminal's speaker.

[0072] Example 1

[0073] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0074] Conventional communication support systems are limited in the processing of text entered by users, making it difficult to quickly and accurately respond to diverse needs such as translation, summarization, and suggesting creative ideas. In addition, the series of processes required to provide generated responses as voice data are not integrated, leaving a need for an improved user experience.

[0075] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0076] In this invention, the server includes means for a user to input text, means for receiving the text and formatting it into a standard data format, means for transmitting the formatted data to the server, means for the server to analyze the received data and request processing from a generative AI model, means for the generative AI model to generate an appropriate response based on the text, means for receiving the generated response and transmitting it to a terminal, means for converting the received response into voice data, and means for providing the voice data to the user. This allows users to quickly and accurately respond to a variety of needs, such as translation, summarization, and creative idea suggestions.

[0077] A "user" is an individual or organization that inputs text and uses the system's functionality based on that input.

[0078] "Text" is a string of characters that a user inputs into the system.

[0079] A "terminal" is a device that allows a user to input text and processes it, such as receiving, sending, and formatting.

[0080] A "standard data format" is a common data format used to format text data, such as JSON.

[0081] A "server" is a computer system that receives data sent from a terminal, analyzes it, and requests processing from the generative AI model.

[0082] A "generative AI model" is an artificial intelligence model that translates text, summarizes it, suggests creative ideas, and more, based on pre-trained data.

[0083] "Means for requesting processing" refers to the method or technology by which the server sends an analysis request to the generated AI model.

[0084] A "response" is the textual result generated by a generative AI model.

[0085] "Audio data" is data obtained by converting text into audio.

[0086] "Speech synthesis" is a technology that converts text into audio data. For example, it uses a speech synthesis API.

[0087] "Means for providing" refers to a method or technology for delivering the generated voice data to the user.

[0088] This invention relates to a communication support system that processes user-entered text using a generative AI model, translating, summarizing, and proposing creative ideas, and then providing the results as voice data. This system is comprised of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0089] A user inputs text through their terminal. The terminal has a text input field, and when the user inputs text, the text is received by the terminal. For example, the user inputs "Translate 'Hello' to Japanese."

[0090] The terminal formats the received text into a standard data format such as JSON. For example, the input text is formatted into JSON as follows:

[0091] json

[0092] {

[0093] "text": "Translate 'Hello' to Japanese"

[0094] }

[0095] The formatted data is sent to a server, which analyzes the received data and determines whether it corresponds to a translation task, a summarization task, or a creative idea proposal task.

[0096] Once the analysis is complete, the server sends an analysis request to the generative AI model. The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. For example, in response to the prompt "Translate 'Hello' to Japanese," the generative AI model will generate the translation result "Hello." This generated response text is then reformatted into JSON format and returned to the server.

[0097] The server receives the response text from the generative AI model and sends it to the device. The device then uses speech synthesis technology to convert the received text into audio data. Specifically, the text is converted into audio through a speech synthesis API (e.g., Google Text-to-Speech API) and the audio data is provided to the user through a speaker. For example, if a user types "This is a long document that needs summarizing...", the generative AI model summarizes the long text and generates summary text that reads "This document needs summarizing." The device then receives this summary text and plays it back to the user through speech synthesis.

[0098] If a user types "Give me an idea for a science project," the generative AI model will generate creative ideas, such as "Build a small wind turbine using household materials." The device will convert this idea into audio data and play it back to the user.

[0099] This allows users to efficiently meet a variety of needs, such as translation, summarization, and creative idea suggestions.The system of the present invention provides multiple functions in an integrated manner and supports user communication, thereby resolving conventional problems.

[0100] The following are examples of prompt sentences:

[0101] 1. "Translate 'Hello' to Japanese" (translation task)

[0102] 2. "Summarize this document: This is a long document that needs summarizing..." (summarization task)

[0103] 3. "Give me an idea for a science project" (creative idea proposal task)

[0104] This allows users to quickly and accurately meet their various language and content needs.

[0105] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0106] Step 1:

[0107] The user enters text into the terminal

[0108] The user enters text into the device's text input field, for example, "Translate 'Hello' to Japanese." The entered text is stored in the device's memory.

[0109] Input: User-entered text "Translate 'Hello' to Japanese"

[0110] Output: Text "Translate 'Hello' to Japanese" stored in the device's memory

[0111] Step 2:

[0112] The device receives the text and formats it into JSON.

[0113] The terminal receives the input text and formats it into a standard data format (e.g., JSON), which ensures data consistency and streamlines subsequent processing.

[0114] Input: User-entered text "Translate 'Hello' to Japanese"

[0115] Output: JSON formatted data { "text": "Translate 'Hello' to Japanese"}

[0116] Step 3:

[0117] The terminal sends the formatted data to the server

[0118] The device sends data formatted in JSON to the server using a communication protocol such as an HTTP POST request.

[0119] Input: JSON format data { "text": "Translate 'Hello' to Japanese"}

[0120] Output: Data sent to the server { "text": "Translate 'Hello' to Japanese"}

[0121] Step 4:

[0122] The server analyzes the data and requests processing from the generative AI model.

[0123] The server analyzes the received data and determines that it is a translation task, then sends an analysis request to the generative AI model.

[0124] Input: Data received by the server { "text": "Translate 'Hello' to Japanese"}

[0125] Output: Parsing request sent to the generative AI model { "task": "translate", "text": "Hello", "target_language": "Japanese"}

[0126] Step 5:

[0127] A generative AI model generates response text

[0128] The generative AI model receives a request from the server and generates the Japanese translation of "Hello," which is "Konnichiwa." The generated response text is then reformatted into JSON and sent back to the server.

[0129] Input: Parsing request sent to the generative AI model { "task": "translate", "text": "Hello", "target_language": "Japanese"}

[0130] Output: Response text sent back to the server from the generative AI model { "translated_text": "Hello"}

[0131] Step 6:

[0132] The server receives the response text and sends it to the device.

[0133] The server sends the response text received from the generative AI model to the terminal.

[0134] Input: Response text returned by the generative AI model { "translated_text": "Hello"}

[0135] Output: Response text sent to the device { "translated_text": "Hello"}

[0136] Step 7:

[0137] The device converts the text into audio data and provides it to the user.

[0138] The device converts the received response text into voice data, and uses a speech synthesis API to convert "hello" into voice format and play it back to the user through the speaker.

[0139] Input: Response text received by the device { "translated_text": "Hello"}

[0140] Output: Voice data "Hello"

[0141] This allows users to efficiently receive text translations, summaries, and creative idea suggestions. For example, if a user types "This is a long document that needs summarizing...", the generative AI model will summarize the content and the device will provide the summary to the user via voice.

[0142] (Application example 1)

[0143] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0144] Conventional factory robots have difficulty understanding instructions and immediately providing specific work procedures and a summary of the situation when they receive them. Furthermore, as their use in multilingual environments increases, translation functions are often inadequate. This can lead to reduced work efficiency and communication problems.

[0145] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0146] In this invention, the server includes: a means for a user to input text; a means for receiving the text and transferring it to a generative AI model for analysis; a means for the generative AI model to generate an appropriate response based on the text; a means for converting the generated response into speech; and a means for providing the speech to the user. The system is installed in a factory robot, and the system generates a summary of work procedures and work status using the generative AI model in response to instructions from a worker and provides the summary by speech. This enables quick understanding of work instructions and status, and maintains high work efficiency even in a multilingual environment.

[0147] "User" refers to the worker who operates the factory robot and issues work instructions and requests information.

[0148] "Means for inputting text" refers to an interface that allows workers to input work instructions and information requests to the robot.

[0149] "Generative AI model" refers to an artificial intelligence model used to analyze input text and generate an appropriate response.

[0150] "Means for transmitting to the generative AI model for analysis" refers to the communications means for transmitting input text to the generative AI model for analysis.

[0151] "Means for generating an appropriate response" refers to the means by which a generative AI model analyzes text and translates, summarizes, or suggests work steps based on a specified task.

[0152] "Means for converting to speech" refers to speech synthesis technology used to convert text generated by a generative AI model into speech data.

[0153] "Means for providing audio to the user" refers to means for letting the worker hear the converted audio data through a speaker or the like.

[0154] The term "system installed on a factory robot" refers to a system including the above-mentioned means implemented on a factory robot, which enables interaction with workers.

[0155] "Work instructions" refer to instructions that specify the specific operations and tasks that workers will perform on factory robots.

[0156] "Summary of work procedures and status" refers to a concise explanation of the current work status and the next operation that a factory robot creates using a generative AI model based on instructions from a worker.

[0157] The system for implementing this invention is installed on a factory robot and utilizes a generative AI model in response to work instructions and information requests from workers. This system is configured as follows.

[0158] First, the user, a worker, inputs work instructions or information requests through a text input interface. The input text is received by the terminal, formatted into JSON format for analysis, and then sent to the server. The text input interface uses an HMI (Human Machine Interface) equipped with a touch panel and keyboard.

[0159] The server analyzes the received text and sends a request to the generative AI model based on its content. This generative AI model is capable of processing text in various languages ​​and generates highly accurate responses based on the Transformer architecture. Depending on the analysis, it translates, summarizes, or suggests work procedures. For example, if a prompt such as "Please tell me the steps required for the next process" is input, the generative AI model will generate the following response: "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."

[0160] The generated response is sent back to the server, which then sends it to the terminal. The terminal uses voice synthesis technology to convert the received text data into voice data. This voice synthesis technology uses gTTS (Google Text-to-Speech), and the generated voice data is provided to the worker through a speaker.

[0161] As a specific example, consider the case where a worker inputs the following text:

[0162] Please tell me the steps I need to take in the next step.

[0163] When this prompt is sent to the generative AI model, specific work procedures like those described above are generated and provided as voice data via the terminal.

[0164] This allows workers to receive the necessary information via voice at the appropriate time, enabling them to work efficiently.In addition, the system can translate and summarize with high accuracy in multilingual environments, enabling smooth communication between workers who speak different languages.

[0165] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0166] Step 1:

[0167] The user inputs text. The user uses a text input interface (such as a touch panel or keyboard) to input work instructions or information requests in text format into the terminal. For example, the user might input a prompt such as, "Please tell me the steps required for the next process."

[0168] Step 2:

[0169] The terminal receives the text and formats it into JSON format. The text entered by the user is received by the terminal and converted into a JSON format suitable for analysis. Specifically, the input text is processed into JSON data represented as key-value pairs. After this processing, the input text is formatted as "{"input": "Please tell me the steps required for the next step"}".

[0170] Step 3:

[0171] The device sends the formatted data to the server. The device then sends the generated JSON data to the server as an HTTP request, including appropriate header information (e.g., Content-Type: application / json).

[0172] Step 4:

[0173] The server analyzes the data and sends a request to the generative AI model. The server analyzes the received JSON data and determines the tasks required for the generative AI model. It then sends a request to the generative AI model API including the determined tasks (in this case, a proposed work procedure). Specifically, the server authenticates using an API key and sends a prompt to the generative AI model.

[0174] Step 5:

[0175] The generative AI model generates a response. The generative AI model analyzes the prompt received from the server and generates an appropriate response. For example, in response to the input, "Please tell me the steps required for the next process," it generates the response, "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."

[0176] Step 6:

[0177] The server receives the generated response and sends it to the terminal. After receiving the response text generated by the generative AI model, the server formats this data again into JSON format and sends it back to the terminal. The data to be sent is in the format "{"response": "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."}".

[0178] Step 7:

[0179] The device converts the response text into speech. The device passes the received response text to a speech synthesis API (e.g., gTTS) and converts it into audio data. The speech synthesis API analyzes the string and generates corresponding audio data (e.g., an mp3 file). The converted audio data is saved on the device.

[0180] Step 8:

[0181] The terminal provides the voice data to the user. The converted voice data is played back through the terminal's speaker, allowing the worker to receive a voice response. Specifically, the terminal plays back the generated voice data file and verbally tells the user the work procedure: "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."

[0182] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0183] This invention relates to a communication support system that processes user-entered text using a generative AI model and an emotion engine, translating, summarizing, suggesting creative ideas, and responding to the user's emotions, providing the results as voice data. This system is composed of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0184] A user inputs text through their device. The device has a text input field, and when the user inputs text, the text is received by the device. This text is then formatted into a standard data format such as JSON and sent to the server.

[0185] The server analyzes the data received from the device and determines whether it includes a translation task, a summary task, a creative idea suggestion, or sentiment analysis. The generative AI model and the emotion engine work together to execute a specific task. Specifically, the server sends an analysis request to the generative AI model to request the necessary processing. The emotion engine also analyzes the user's sentiment from the input text, and the generative AI model adjusts its response accordingly.

[0186] The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. The emotion engine can recognize the user's emotions from a variety of input data, including text, voice, and facial expressions. For example, if the text entered by the user contains the emotion "I'm feeling sad today," the emotion engine will identify the emotion "sad," and the generative AI model will generate an appropriate response based on this.

[0187] The response text generated by the generative AI model is formatted in JSON and returned to the server. The server receives this generated response text and sends it back to the device. The device then uses speech synthesis technology to convert the received text into audio data. The text is converted into audio through a speech synthesis API, and the audio data is provided to the user through the speaker.

[0188] For example, if a user types "Translate 'Hello' to Japanese," the generative AI model will generate the translation "Hello." This translation result is received by the device and played back to the user via speech synthesis.

[0189] Also, if a user types, "I had a bad day today," the emotion engine recognizes the emotion "negative," and the generative AI model generates an appropriate response, such as a message like, "I'm sorry to hear that. Would you like to talk about it?"

[0190] Furthermore, if a user types "Give me an idea for a science project," the generative AI model will generate the creative idea "Build a small wind turbine using household materials," which is also provided to the user via voice synthesis.

[0191] In this way, the system of the present invention allows users to efficiently translate, summarize, suggest creative ideas, and even receive emotional responses, thereby solving the problems of the past.

[0192] The processing flow will be explained below.

[0193] Step 1:

[0194] A user enters text into an input field on a terminal, for example, "I had a bad day today."

[0195] Step 2:

[0196] The device captures the user's input text and formats it into JSON, e.g., {"text": "I had a bad day today"}.

[0197] Step 3:

[0198] The device sends the formatted data to the server, using an Internet connection to send a request to the server.

[0199] Step 4:

[0200] The server analyzes the data received from the device. It parses the received JSON data and extracts the text portion.

[0201] Step 5:

[0202] The server transfers the analyzed data to the emotion engine. Specifically, it calls the emotion engine's API and requests emotion analysis. For example, make the following API call:

[0203] python

[0204] emotion_result = emotion_engine.analyze(text="I had a bad day today")

[0205] Step 6:

[0206] The emotion engine analyzes the text "I had a bad day today" and recognizes the emotion "negative." This emotion information is returned to the server.

[0207] Step 7:

[0208] The server transfers the emotion information received from the emotion engine to the generative AI model. At the same time, it also sends the original text. For example, make the following API call:

[0209] python

[0210] response = ai_model.generate_response(text="I had a bad day today", emotion="negative")

[0211] Step 8:

[0212] A generative AI model generates an appropriate response based on the text and sentiment information, e.g., "I'm sorry to hear that. Would you like to talk about it?"

[0213] Step 9:

[0214] The server formats the response received from the generated AI model into JSON format. For example, it generates data like {"response_text": "I'm sorry to hear that. Would you like to talk about it?"}.

[0215] Step 10:

[0216] The server then sends the formatted data to the device, again using the internet connection to return a response to the device.

[0217] Step 11:

[0218] The device parses the JSON data received from the server, interpreting the format {"response_text": "I'm sorry to hear that. Would you like to talk about it?"} and extracts the text portion.

[0219] Step 12:

[0220] Convert text received by the device into speech. Use a speech synthesis API to convert text into speech data. Example: Use the following API:

[0221] python

[0222] synthesized_audio = text_to_speech("I'm sorry to hear that. Would you like to talk about it?")

[0223] Step 13:

[0224] The terminal plays the generated voice data to the user, and provides the voice to the user through the terminal's speaker.

[0225] This allows the user to receive an appropriate voice response that corresponds to their own emotions.

[0226] Example 2

[0227] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0228] Conventional communication support systems have a limited ability to properly analyze text entered by users and provide responses based on that content. Furthermore, they are unable to generate responses that reflect the user's emotions, making it difficult to improve the quality of communication. Furthermore, functions such as text translation and summarization are often provided separately, making them less convenient as an integrated system.

[0229] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for a user to input text, means for receiving the text, shaping it into a data format, and then transmitting it to the server for analysis, means for the server to analyze the text and determine a task, means for the generative AI model and emotion engine to perform appropriate processing based on the task, means for converting the generated response into speech, and means for providing the speech to the user. This makes it possible to automatically provide appropriate translations, summaries, suggestions for creative ideas, and even responses according to emotions based on the text entered by the user.

[0230] "User" refers to any individual or corporation that uses this system.

[0231] "Means for inputting text" refers to an interface for a user to input text information, such as a keyboard or a touch screen.

[0232] "Data format" refers to a format for formatting information based on certain rules, and includes, for example, JSON format.

[0233] A "server" refers to a computer system that provides services to other computers over a network.

[0234] A "generative AI model" refers to an artificial intelligence model that is trained using large amounts of text data and generates and analyzes text.

[0235] "Emotion engine" refers to a system for analyzing a user's emotions from input text.

[0236] "Means for converting a response into voice" refers to technology for converting text data into voice data, including, for example, a voice synthesis API.

[0237] A "task" refers to a specific process or function that the system must perform, such as translation, summarization, creative idea suggestion, or sentiment analysis.

[0238] "Analysis" refers to the process of analyzing input text data and clarifying its content and intent.

[0239] This invention relates to a communication support system that performs advanced processing of text entered by a user, and provides translation, summarization, creative idea suggestions, and even responses tailored to the user's emotions. This system consists of three main components: the user, the terminal, and the server. Each component and its operation are described in detail below.

[0240] User operations and device roles

[0241] A user inputs text through a terminal such as a personal computer or smartphone. The terminal is provided with a text input field, and when the user inputs text, the text is received by the terminal. The input text is formatted into a standard data format such as JSON. This formatted data is then sent to the server via an HTTP request.

[0242] Data analysis and model execution on the server

[0243] The server analyzes the received data and determines whether it corresponds to a translation task, a summarization task, a creative idea suggestion, or sentiment analysis. This determination process uses a text analysis library (e.g., NLTK or spaCy). After determining the specific task, the server invokes a generative AI model and sentiment engine to perform the necessary processing.

[0244] Generative AI models are pre-trained using large amounts of text data, enabling accurate analysis and response generation. For example, OpenAI's GPT-3 is used. Meanwhile, emotion engines are systems that recognize user emotions from input text, such as IBM Watson's Tone Analyzer.

[0245] Response generation and information return

[0246] The response text generated by the generative AI model is again formatted in JSON format. This response data is sent from the server to the device. The device then uses speech synthesis technology to convert the received response text into audio data. For example, Google Cloud Text-to-Speech API is used. The converted audio data is then provided to the user via the speaker.

[0247] Specific examples

[0248] Below is a concrete example of a prompt entered by a user and the system's response to it.

[0249] Example 1: Translation task

[0250] When a user types "Translate 'Hello' to Japanese," the device sends the text to the server, and the generative AI model generates the translation result "Hello," which is then converted into speech on the device and played back to the user.

[0251] Example 2: Emotional response task

[0252] If a user types "I had a bad day today," the emotion engine recognizes the emotion "negative," and the generative AI model generates the corresponding message, "I'm sorry to hear that. Would you like to talk about it?" This message is also converted to audio and provided to the user.

[0253] Example 3: Creative idea proposal task

[0254] If a user types "Give me an idea for a science project," the generative AI model generates the idea "Build a small wind turbine using household materials," which is also provided to the user via voice synthesis.

[0255] In this way, each component works in cooperation with the others to efficiently and flexibly provide a variety of responses according to the text entered by the user.

[0256] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0257] Program processing flow and specific explanation

[0258] Step 1:

[0259] The user inputs text. For example, the user inputs "Translate 'Hello' to Japanese" into a text input field on the terminal. The input text is received by the terminal.

[0260] Step 2:

[0261] The device formats the received text into JSON format. Specifically, it uses JavaScript to convert the input text into an object format, and then uses the JSON.stringify function to convert it into a JSON string. This formatted data is sent as an HTTP POST request to the API endpoint.

[0262] Input: Raw text entered by the user ("Translate 'Hello' to Japanese")

[0263] Output: JSON formatted data

[0264] Step 3:

[0265] The server receives the request and parses the received JSON data. The server receives the request using the Python Flask framework, and parses the JSON data using request.get_json().

[0266] Input: JSON format data

[0267] Output: Parsed text ("Translate 'Hello' to Japanese")

[0268] Step 4:

[0269] The server analyzes the parsed text to determine whether it is a translation, summary, creative idea suggestion, or sentiment analysis, for example using a text analysis library (NLTK or spaCy).

[0270] Input: Parsed text

[0271] Output: The type of task (in this case, a translation task)

[0272] Step 5:

[0273] The server calls a generative AI model or emotion engine based on the identified task. For example, for a translation task, it sends a translation request to a generative AI model (e.g., GPT-3). In this case, it uses the GPT-3 API to send a request to translate "Hello" into Japanese.

[0274] Input: Task type (translation task) and source text ("Hello")

[0275] Output: Translation result("Hello")

[0276] Step 6:

[0277] The generated response text is then formatted into JSON again. The server then uses the jsonify function to convert the translation result into JSON format and sends it back as an HTTP response.

[0278] Input: Translation result("Hello")

[0279] Output: Response data in JSON format

[0280] Step 7:

[0281] The device analyzes the received response data and converts it into audio. The device uses JavaScript to parse the JSON data and converts the translation results into audio data using a speech synthesis API (such as Google Cloud Text-to-Speech). The generated audio data is played using the browser's Audio object.

[0282] Input: Response data in JSON format

[0283] Output: Audio data

[0284] Step 8:

[0285] The device plays the generated audio data through the speaker and provides it to the user as sound. Specifically, it plays the audio using the Audio.play() method.

[0286] Input: Audio data

[0287] Output: Speech provided to the user ("Hello")

[0288] The above is the specific processing flow of this system and the detailed operation of each step, which allows users to efficiently and flexibly obtain results for translation and other tasks.

[0289] (Application example 2)

[0290] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0291] Conventional food delivery services require users to input their menu items through unintuitive text input, and are unable to provide appropriate suggestions based on the user's emotions. This results in a poor user experience and reduces the appeal of the service.

[0292] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input text, means for receiving the text and transferring it to a generative AI model for analysis, means for the generative AI model to generate an appropriate response based on the text, means for converting the generated response into speech, means for providing the speech to the user, means for analyzing the user's emotions using the generative AI model and an emotion engine, and means for suggesting an appropriate menu to the user based on the emotion analysis. This allows the user to intuitively place an order using natural language and receive appropriate menu suggestions based on their emotions.

[0293] A "means for user text input" is a device or software that provides an interface for a user to input text in a natural language.

[0294] The "means for receiving said text and forwarding it to a generative AI model for analysis" refers to a function that sends user-entered text to a server and prepares the data to be analyzed by a generative AI model.

[0295] The "means by which the generative AI model generates an appropriate response based on the text" refers to the algorithm or process by which the generative AI model generates an appropriate reply or suggestion based on the text analyzed by the generative AI model.

[0296] "Means for converting the generated response into speech" refers to a technology or system that converts the text response generated by the generative AI model into speech data.

[0297] The "means for providing the audio to the user" refers to a device such as a speaker or earphone for playing back the data converted into audio to the user.

[0298] The "means for analyzing user emotions using the generative AI model and emotion engine" refers to a system that combines a generative AI model and an emotion analysis engine to identify emotions from the user's input text and adjust responses and suggestions based on the results.

[0299] The "means for suggesting an appropriate menu to a user based on the emotion analysis" is a system that automatically suggests an appropriate menu to a user in a food delivery service based on the user's emotions analyzed by an emotion analysis engine.

[0300] This invention is a system that processes user-entered text using a generative AI model and an emotion engine to support ordering and suggestions in food delivery services. This system consists of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0301] A user inputs text through their device. The device has a text input field, and when the user types, the text is received by the device. This text is then formatted into a standard data format such as JSON and sent to the server.

[0302] The server analyzes the data received from the device and determines whether it includes a translation task, a summary task, creative suggestions, or sentiment analysis. The generative AI model and the emotion engine work together to execute a specific task. Specifically, the server sends an analysis request to the generative AI model to request the necessary processing. The emotion engine also analyzes the user's sentiment from the input text, and the generative AI model adjusts its response accordingly.

[0303] The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. The emotion engine can recognize the user's emotions from a variety of input data, including text, voice, and facial expressions. For example, if the text entered by the user contains emotions such as "I had a tough day, recommend me something soothing," the emotion engine will identify emotions such as fatigue and anxiety, and the generative AI model will generate an appropriate response based on this.

[0304] The response text generated by the generative AI model is formatted in JSON and returned to the server. The server receives this generated response text and sends it back to the device. The device then uses speech synthesis technology to convert the received text into audio data. The text is converted into audio through a speech synthesis API, and the audio data is provided to the user through the speaker.

[0305] For example, if a user types "I want to order something spicy," the generative AI model will suggest "spicy dish suggestions," which will be received by the device and played aloud to the user via speech synthesis.

[0306] For example, if a user types, "I had a tough day, recommend me something soothing," the emotion engine will recognize emotions such as "fatigue" and "anxiety," and the generative AI model will generate an appropriate response, such as "recommended smoothies to relax you."

[0307] In this way, the system of the present invention allows users to efficiently obtain translations, summaries, and creative suggestions, as well as receive appropriate menu suggestions based on their current emotions, thereby improving the user experience in food delivery services and increasing their appeal.

[0308] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0309] Step 1:

[0310] The user inputs text into the terminal. The input text includes the user's order details or questions. For example, "I want to order something spicy." The input is formatted into a standard data format such as JSON. The input data is sent from the terminal to the server.

[0311] Step 2:

[0312] The server receives the text sent from the device, stores it in a database, and formats it appropriately for analysis. The server analyzes the text and determines whether it needs to be translated, summarized, suggested, or sentiment analyzed.

[0313] Step 3:

[0314] The server sends an analysis request to the generative AI model, which generates a highly accurate response based on the received text. For example, if a user types "I want to order something spicy," the generative AI model generates "spicy dish suggestions."

[0315] Step 4:

[0316] The server receives the response from the generative AI model and passes the text to the emotion engine to analyze the user's emotions. The emotion engine identifies the emotions contained in the text and adjusts the response accordingly. For example, for the input "I had a tough day, recommend me something soothing," the emotion engine identifies emotions such as "fatigue" and "anxiety."

[0317] Step 5:

[0318] The server combines the generated response text with the sentiment analysis results to generate optimal suggestions and responses, and the combined results are formatted in JSON format.

[0319] Step 6:

[0320] The server sends the formatted response text to the device, which then uses speech synthesis technology to provide a response to the user. The speech synthesis API converts the text into speech and provides it to the user as audio data through the speaker.

[0321] Step 7:

[0322] Users receive responses and suggestions through voice data, such as suggestions for spicy dishes or relaxing menu items, based on their emotions and the order they make.

[0323] In this way, a system is constructed that uses a generative AI model and emotion engine to perform advanced analysis and data calculations based on the user's input at each processing step, providing the user with an appropriate response.

[0324] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0325] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0326] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0327] [Second embodiment]

[0328] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0329] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0330] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0331] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0332] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0333] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0334] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0335] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0336] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0337] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0338] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0339] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0340] This invention relates to a communication support system that processes user-entered text using a generative AI model, translating, summarizing, and proposing creative ideas, and then providing the results as voice data. This system is comprised of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0341] The user inputs text through their device. The device has a text input field, and when the user inputs text, the text is received by the device. The received text is formatted into a standard data format such as JSON and then sent to the server.

[0342] The server analyzes the data received from the device and determines whether it corresponds to a translation task, a summarization task, or a creative idea suggestion. The generative AI model is then used for this processing. Specifically, the server sends an analysis request to the generative AI model, which then translates, summarizes, or generates creative ideas as needed.

[0343] The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. For example, when a user enters the text "Translate 'Hello' to Japanese," the generative AI model generates the translation result "Hello." This generated response text is then reformatted into JSON format and returned to the server.

[0344] The server receives the response text from the generative AI model and sends it to the device. The device uses speech synthesis technology to convert the received text into audio data. The text is converted into audio through a speech synthesis API, and the audio data is provided to the user through the speaker.

[0345] For example, if a user types "This is a long document that needs summarizing...", the generative AI model will summarize the long sentence and generate the summary text "This document needs summarizing". The device will receive this summary text and play it back to the user via speech synthesis.

[0346] If a user types "Give me an idea for a science project," the generative AI model will generate creative ideas, such as "Build a small wind turbine using household materials." The device will convert this idea into audio data and play it back to the user.

[0347] This allows users to efficiently meet a variety of needs, such as translation, summarization, and creative idea suggestions.The system of the present invention provides multiple functions in an integrated manner to support user communication, thereby resolving conventional problems.

[0348] The processing flow will be explained below.

[0349] Step 1:

[0350] A user enters text into an input field on a terminal. For example, "Translate 'Hello' to Japanese."

[0351] Step 2:

[0352] The device captures the user's input text and formats it into JSON format. For example, it generates data like {"action": "translate", "text": "Hello", "target_language": "Japanese"}.

[0353] Step 3:

[0354] The device sends the formatted data to the server, using an Internet connection to send a request to the server.

[0355] Step 4:

[0356] The server analyzes the data received from the device. It parses the received JSON data and checks that the "action" field is "translate".

[0357] Step 5:

[0358] The server transfers the analyzed data to the generative AI model. Specifically, it calls the generative AI model's API and requests a translation task. For example, make the following API call:

[0359] python

[0360] translation_result = ai_model.translate(text="Hello", target_language="Japanese")

[0361] Step 6:

[0362] The generative AI model translates the text "Hello" to "Hello." The translation result is returned to the server.

[0363] Step 7:

[0364] The server formats the translation results received from the AI ​​model into JSON format. For example, it generates data like {"translated_text": "Hello"}.

[0365] Step 8:

[0366] The server sends the formatted data to the device, again using the internet connection to return a response to the device.

[0367] Step 9:

[0368] The device parses the JSON data received from the server, interpreting the format {"translated_text": "Hello"} and extracting the contents of the "translated_text" field.

[0369] Step 10:

[0370] Convert text received by the device into speech. Use a speech synthesis API to convert text into speech data. Example: Use the following API:

[0371] python

[0372] synthesized_audio = text_to_speech("Hello")

[0373] Step 11:

[0374] The terminal plays the generated voice data to the user, and provides the voice to the user through the terminal's speaker.

[0375] Example 1

[0376] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0377] Conventional communication support systems are limited in the processing of text entered by users, making it difficult to quickly and accurately respond to diverse needs such as translation, summarization, and suggesting creative ideas. In addition, the series of processes required to provide generated responses as voice data are not integrated, leaving a need for an improved user experience.

[0378] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0379] In this invention, the server includes means for a user to input text, means for receiving the text and formatting it into a standard data format, means for transmitting the formatted data to the server, means for the server to analyze the received data and request processing from a generative AI model, means for the generative AI model to generate an appropriate response based on the text, means for receiving the generated response and transmitting it to a terminal, means for converting the received response into voice data, and means for providing the voice data to the user. This allows users to quickly and accurately respond to a variety of needs, such as translation, summarization, and creative idea suggestions.

[0380] A "user" is an individual or organization that inputs text and uses the system's functionality based on that input.

[0381] "Text" is a string of characters that a user inputs into the system.

[0382] A "terminal" is a device that allows a user to input text and processes it, such as receiving, sending, and formatting.

[0383] A "standard data format" is a common data format used to format text data, such as JSON.

[0384] A "server" is a computer system that receives data sent from a terminal, analyzes it, and requests processing from the generative AI model.

[0385] A "generative AI model" is an artificial intelligence model that translates text, summarizes it, suggests creative ideas, and more, based on pre-trained data.

[0386] "Means for requesting processing" refers to the method or technology by which the server sends an analysis request to the generated AI model.

[0387] A "response" is the textual result generated by a generative AI model.

[0388] "Audio data" is data obtained by converting text into audio.

[0389] "Speech synthesis" is a technology that converts text into audio data. For example, it uses a speech synthesis API.

[0390] "Means for providing" refers to a method or technology for delivering the generated voice data to the user.

[0391] This invention relates to a communication support system that processes user-entered text using a generative AI model, translating, summarizing, and proposing creative ideas, and then providing the results as voice data. This system is comprised of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0392] A user inputs text through their terminal. The terminal has a text input field, and when the user inputs text, the text is received by the terminal. For example, the user inputs "Translate 'Hello' to Japanese."

[0393] The terminal formats the received text into a standard data format such as JSON. For example, the input text is formatted into JSON as follows:

[0394] json

[0395] {

[0396] "text": "Translate 'Hello' to Japanese"

[0397] }

[0398] The formatted data is sent to a server, which analyzes the received data and determines whether it corresponds to a translation task, a summarization task, or a creative idea proposal task.

[0399] Once the analysis is complete, the server sends an analysis request to the generative AI model. The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. For example, in response to the prompt "Translate 'Hello' to Japanese," the generative AI model will generate the translation result "Hello." This generated response text is then reformatted into JSON format and returned to the server.

[0400] The server receives the response text from the generative AI model and sends it to the device. The device then uses speech synthesis technology to convert the received text into audio data. Specifically, the text is converted into audio through a speech synthesis API (e.g., Google Text-to-Speech API) and the audio data is provided to the user through a speaker. For example, if a user types "This is a long document that needs summarizing...", the generative AI model summarizes the long text and generates summary text that reads "This document needs summarizing." The device then receives this summary text and plays it back to the user through speech synthesis.

[0401] If a user types "Give me an idea for a science project," the generative AI model will generate creative ideas, such as "Build a small wind turbine using household materials." The device will convert this idea into audio data and play it back to the user.

[0402] This allows users to efficiently meet a variety of needs, such as translation, summarization, and creative idea suggestions.The system of the present invention provides multiple functions in an integrated manner and supports user communication, thereby resolving conventional problems.

[0403] The following are examples of prompt sentences:

[0404] 1. "Translate 'Hello' to Japanese" (translation task)

[0405] 2. "Summarize this document: This is a long document that needs summarizing..." (summarization task)

[0406] 3. "Give me an idea for a science project" (creative idea proposal task)

[0407] This allows users to quickly and accurately meet their various language and content needs.

[0408] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0409] Step 1:

[0410] The user enters text into the terminal

[0411] The user enters text into the device's text input field, for example, "Translate 'Hello' to Japanese." The entered text is stored in the device's memory.

[0412] Input: User-entered text "Translate 'Hello' to Japanese"

[0413] Output: Text "Translate 'Hello' to Japanese" stored in the device's memory

[0414] Step 2:

[0415] The device receives the text and formats it into JSON.

[0416] The terminal receives the input text and formats it into a standard data format (e.g., JSON), which ensures data consistency and streamlines subsequent processing.

[0417] Input: User-entered text "Translate 'Hello' to Japanese"

[0418] Output: JSON formatted data { "text": "Translate 'Hello' to Japanese"}

[0419] Step 3:

[0420] The terminal sends the formatted data to the server

[0421] The device sends data formatted in JSON to the server using a communication protocol such as an HTTP POST request.

[0422] Input: JSON format data { "text": "Translate 'Hello' to Japanese"}

[0423] Output: Data sent to the server { "text": "Translate 'Hello' to Japanese"}

[0424] Step 4:

[0425] The server analyzes the data and requests processing from the generative AI model.

[0426] The server analyzes the received data and determines that it is a translation task, then sends an analysis request to the generative AI model.

[0427] Input: Data received by the server { "text": "Translate 'Hello' to Japanese"}

[0428] Output: Parsing request sent to the generative AI model { "task": "translate", "text": "Hello", "target_language": "Japanese"}

[0429] Step 5:

[0430] A generative AI model generates response text

[0431] The generative AI model receives a request from the server and generates the Japanese translation of "Hello," which is "Konnichiwa." The generated response text is then reformatted into JSON and sent back to the server.

[0432] Input: Parsing request sent to the generative AI model { "task": "translate", "text": "Hello", "target_language": "Japanese"}

[0433] Output: Response text sent back to the server from the generative AI model { "translated_text": "Hello"}

[0434] Step 6:

[0435] The server receives the response text and sends it to the device.

[0436] The server sends the response text received from the generative AI model to the terminal.

[0437] Input: Response text returned by the generative AI model { "translated_text": "Hello"}

[0438] Output: Response text sent to the device { "translated_text": "Hello"}

[0439] Step 7:

[0440] The device converts the text into audio data and provides it to the user.

[0441] The device converts the received response text into voice data, and uses a speech synthesis API to convert "hello" into voice format and play it back to the user through the speaker.

[0442] Input: Response text received by the device { "translated_text": "Hello"}

[0443] Output: Voice data "Hello"

[0444] This allows users to efficiently receive text translations, summaries, and creative idea suggestions. For example, if a user types "This is a long document that needs summarizing...", the generative AI model will summarize the content and the device will provide the summary to the user via voice.

[0445] (Application example 1)

[0446] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0447] Conventional factory robots have difficulty understanding instructions and immediately providing specific work procedures and a summary of the situation when they receive them. Furthermore, as their use in multilingual environments increases, translation functions are often inadequate. This can lead to reduced work efficiency and communication problems.

[0448] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0449] In this invention, the server includes: a means for a user to input text; a means for receiving the text and transferring it to a generative AI model for analysis; a means for the generative AI model to generate an appropriate response based on the text; a means for converting the generated response into speech; and a means for providing the speech to the user. The system is installed in a factory robot, and the system generates a summary of work procedures and work status using the generative AI model in response to instructions from a worker and provides the summary by speech. This enables quick understanding of work instructions and status, and maintains high work efficiency even in a multilingual environment.

[0450] "User" refers to the worker who operates the factory robot and issues work instructions and requests information.

[0451] "Means for inputting text" refers to an interface that allows workers to input work instructions and information requests to the robot.

[0452] "Generative AI model" refers to an artificial intelligence model used to analyze input text and generate an appropriate response.

[0453] "Means for transmitting to the generative AI model for analysis" refers to the communications means for transmitting input text to the generative AI model for analysis.

[0454] "Means for generating an appropriate response" refers to the means by which a generative AI model analyzes text and translates, summarizes, or suggests work steps based on a specified task.

[0455] "Means for converting to speech" refers to speech synthesis technology used to convert text generated by a generative AI model into speech data.

[0456] "Means for providing audio to the user" refers to means for letting the worker hear the converted audio data through a speaker or the like.

[0457] The term "system installed on a factory robot" refers to a system including the above-mentioned means implemented on a factory robot, which enables interaction with workers.

[0458] "Work instructions" refer to instructions that specify the specific operations and tasks that workers will perform on factory robots.

[0459] "Summary of work procedures and status" refers to a concise explanation of the current work status and the next operation that a factory robot creates using a generative AI model based on instructions from a worker.

[0460] The system for implementing this invention is installed on a factory robot and utilizes a generative AI model in response to work instructions and information requests from workers. This system is configured as follows.

[0461] First, the user, a worker, inputs work instructions or information requests through a text input interface. The input text is received by the terminal, formatted into JSON format for analysis, and then sent to the server. The text input interface uses an HMI (Human Machine Interface) equipped with a touch panel and keyboard.

[0462] The server analyzes the received text and sends a request to the generative AI model based on its content. This generative AI model is capable of processing text in various languages ​​and generates highly accurate responses based on the Transformer architecture. Depending on the analysis, it translates, summarizes, or suggests work procedures. For example, if a prompt such as "Please tell me the steps required for the next process" is input, the generative AI model will generate the following response: "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."

[0463] The generated response is sent back to the server, which then sends it to the terminal. The terminal uses voice synthesis technology to convert the received text data into voice data. This voice synthesis technology uses gTTS (Google Text-to-Speech), and the generated voice data is provided to the worker through a speaker.

[0464] As a specific example, consider the case where a worker inputs the following text:

[0465] Please tell me the steps I need to take in the next step.

[0466] When this prompt is sent to the generative AI model, specific work procedures like those described above are generated and provided as voice data via the terminal.

[0467] This allows workers to receive the necessary information via voice at the appropriate time, enabling them to work efficiently.In addition, the system can translate and summarize with high accuracy in multilingual environments, enabling smooth communication between workers who speak different languages.

[0468] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0469] Step 1:

[0470] The user inputs text. The user uses a text input interface (such as a touch panel or keyboard) to input work instructions or information requests in text format into the terminal. For example, the user might input a prompt such as, "Please tell me the steps required for the next process."

[0471] Step 2:

[0472] The terminal receives the text and formats it into JSON format. The text entered by the user is received by the terminal and converted into a JSON format suitable for analysis. Specifically, the input text is processed into JSON data represented as key-value pairs. After this processing, the input text is formatted as "{"input": "Please tell me the steps required for the next step"}".

[0473] Step 3:

[0474] The device sends the formatted data to the server. The device then sends the generated JSON data to the server as an HTTP request, including appropriate header information (e.g., Content-Type: application / json).

[0475] Step 4:

[0476] The server analyzes the data and sends a request to the generative AI model. The server analyzes the received JSON data and determines the tasks required for the generative AI model. It then sends a request to the generative AI model API including the determined tasks (in this case, a proposed work procedure). Specifically, the server authenticates using an API key and sends a prompt to the generative AI model.

[0477] Step 5:

[0478] The generative AI model generates a response. The generative AI model analyzes the prompt received from the server and generates an appropriate response. For example, in response to the input, "Please tell me the steps required for the next process," it generates the response, "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."

[0479] Step 6:

[0480] The server receives the generated response and sends it to the terminal. After receiving the response text generated by the generative AI model, the server formats this data again into JSON format and sends it back to the terminal. The data to be sent is in the format "{"response": "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."}".

[0481] Step 7:

[0482] The device converts the response text into speech. The device passes the received response text to a speech synthesis API (e.g., gTTS) and converts it into audio data. The speech synthesis API analyzes the string and generates corresponding audio data (e.g., an mp3 file). The converted audio data is saved on the device.

[0483] Step 8:

[0484] The terminal provides the voice data to the user. The converted voice data is played back through the terminal's speaker, allowing the worker to receive a voice response. Specifically, the terminal plays back the generated voice data file and verbally tells the user the work procedure: "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."

[0485] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0486] This invention relates to a communication support system that processes user-entered text using a generative AI model and an emotion engine, translating, summarizing, suggesting creative ideas, and responding to the user's emotions, providing the results as voice data. This system is composed of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0487] A user inputs text through their device. The device has a text input field, and when the user inputs text, the text is received by the device. This text is then formatted into a standard data format such as JSON and sent to the server.

[0488] The server analyzes the data received from the device and determines whether it includes a translation task, a summary task, a creative idea suggestion, or sentiment analysis. The generative AI model and the emotion engine work together to execute a specific task. Specifically, the server sends an analysis request to the generative AI model to request the necessary processing. The emotion engine also analyzes the user's sentiment from the input text, and the generative AI model adjusts its response accordingly.

[0489] The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. The emotion engine can recognize the user's emotions from a variety of input data, including text, voice, and facial expressions. For example, if the text entered by the user contains the emotion "I'm feeling sad today," the emotion engine will identify the emotion "sad," and the generative AI model will generate an appropriate response based on this.

[0490] The response text generated by the generative AI model is formatted in JSON and returned to the server. The server receives this generated response text and sends it back to the device. The device then uses speech synthesis technology to convert the received text into audio data. The text is converted into audio through a speech synthesis API, and the audio data is provided to the user through the speaker.

[0491] For example, if a user types "Translate 'Hello' to Japanese," the generative AI model will generate the translation "Hello." This translation result is received by the device and played back to the user via speech synthesis.

[0492] Also, if a user types, "I had a bad day today," the emotion engine recognizes the emotion "negative," and the generative AI model generates an appropriate response, such as a message like, "I'm sorry to hear that. Would you like to talk about it?"

[0493] Furthermore, if a user types "Give me an idea for a science project," the generative AI model will generate the creative idea "Build a small wind turbine using household materials," which is also provided to the user via voice synthesis.

[0494] In this way, the system of the present invention allows users to efficiently translate, summarize, suggest creative ideas, and even receive emotional responses, thereby solving the problems of the past.

[0495] The processing flow will be explained below.

[0496] Step 1:

[0497] A user enters text into an input field on a terminal, for example, "I had a bad day today."

[0498] Step 2:

[0499] The device captures the user's input text and formats it into JSON, e.g., {"text": "I had a bad day today"}.

[0500] Step 3:

[0501] The device sends the formatted data to the server, using an Internet connection to send a request to the server.

[0502] Step 4:

[0503] The server analyzes the data received from the device. It parses the received JSON data and extracts the text portion.

[0504] Step 5:

[0505] The server transfers the analyzed data to the emotion engine. Specifically, it calls the emotion engine's API and requests emotion analysis. For example, make the following API call:

[0506] python

[0507] emotion_result = emotion_engine.analyze(text="I had a bad day today")

[0508] Step 6:

[0509] The emotion engine analyzes the text "I had a bad day today" and recognizes the emotion "negative." This emotion information is returned to the server.

[0510] Step 7:

[0511] The server transfers the emotion information received from the emotion engine to the generative AI model. At the same time, it also sends the original text. For example, make the following API call:

[0512] python

[0513] response = ai_model.generate_response(text="I had a bad day today", emotion="negative")

[0514] Step 8:

[0515] A generative AI model generates an appropriate response based on the text and sentiment information, e.g., "I'm sorry to hear that. Would you like to talk about it?"

[0516] Step 9:

[0517] The server formats the response received from the generated AI model into JSON format. For example, it generates data like {"response_text": "I'm sorry to hear that. Would you like to talk about it?"}.

[0518] Step 10:

[0519] The server then sends the formatted data to the device, again using the internet connection to return a response to the device.

[0520] Step 11:

[0521] The device parses the JSON data received from the server, interpreting the format {"response_text": "I'm sorry to hear that. Would you like to talk about it?"} and extracts the text portion.

[0522] Step 12:

[0523] Convert text received by the device into speech. Use a speech synthesis API to convert text into speech data. Example: Use the following API:

[0524] python

[0525] synthesized_audio = text_to_speech("I'm sorry to hear that. Would you like to talk about it?")

[0526] Step 13:

[0527] The terminal plays the generated voice data to the user, and provides the voice to the user through the terminal's speaker.

[0528] This allows the user to receive an appropriate voice response that corresponds to their own emotions.

[0529] Example 2

[0530] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0531] Conventional communication support systems have a limited ability to properly analyze text entered by users and provide responses based on that content. Furthermore, they are unable to generate responses that reflect the user's emotions, making it difficult to improve the quality of communication. Furthermore, functions such as text translation and summarization are often provided separately, making them less convenient as an integrated system.

[0532] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for a user to input text, means for receiving the text, shaping it into a data format, and then transmitting it to the server for analysis, means for the server to analyze the text and determine a task, means for the generative AI model and emotion engine to perform appropriate processing based on the task, means for converting the generated response into speech, and means for providing the speech to the user. This makes it possible to automatically provide appropriate translations, summaries, suggestions for creative ideas, and even responses according to emotions based on the text entered by the user.

[0533] "User" refers to any individual or corporation that uses this system.

[0534] "Means for inputting text" refers to an interface for a user to input text information, such as a keyboard or a touch screen.

[0535] "Data format" refers to a format for formatting information based on certain rules, and includes, for example, JSON format.

[0536] A "server" refers to a computer system that provides services to other computers over a network.

[0537] A "generative AI model" refers to an artificial intelligence model that is trained using large amounts of text data and generates and analyzes text.

[0538] "Emotion engine" refers to a system for analyzing a user's emotions from input text.

[0539] "Means for converting a response into voice" refers to technology for converting text data into voice data, including, for example, a voice synthesis API.

[0540] A "task" refers to a specific process or function that the system must perform, such as translation, summarization, creative idea suggestion, or sentiment analysis.

[0541] "Analysis" refers to the process of analyzing input text data and clarifying its content and intent.

[0542] This invention relates to a communication support system that performs advanced processing of text entered by a user, and provides translation, summarization, creative idea suggestions, and even responses tailored to the user's emotions. This system consists of three main components: the user, the terminal, and the server. Each component and its operation are described in detail below.

[0543] User operations and device roles

[0544] A user inputs text through a terminal such as a personal computer or smartphone. The terminal is provided with a text input field, and when the user inputs text, the text is received by the terminal. The input text is formatted into a standard data format such as JSON. This formatted data is then sent to the server via an HTTP request.

[0545] Data analysis and model execution on the server

[0546] The server analyzes the received data and determines whether it corresponds to a translation task, a summarization task, a creative idea suggestion, or sentiment analysis. This determination process uses a text analysis library (e.g., NLTK or spaCy). After determining the specific task, the server invokes a generative AI model and sentiment engine to perform the necessary processing.

[0547] Generative AI models are pre-trained using large amounts of text data, enabling accurate analysis and response generation. For example, OpenAI's GPT-3 is used. Meanwhile, emotion engines are systems that recognize user emotions from input text, such as IBM Watson's Tone Analyzer.

[0548] Response generation and information return

[0549] The response text generated by the generative AI model is again formatted in JSON format. This response data is sent from the server to the device. The device then uses speech synthesis technology to convert the received response text into audio data. For example, Google Cloud Text-to-Speech API is used. The converted audio data is then provided to the user via the speaker.

[0550] Specific examples

[0551] Below is a concrete example of a prompt entered by a user and the system's response to it.

[0552] Example 1: Translation task

[0553] When a user types "Translate 'Hello' to Japanese," the device sends the text to the server, and the generative AI model generates the translation result "Hello," which is then converted into speech on the device and played back to the user.

[0554] Example 2: Emotional response task

[0555] If a user types "I had a bad day today," the emotion engine recognizes the emotion "negative," and the generative AI model generates the corresponding message, "I'm sorry to hear that. Would you like to talk about it?" This message is also converted to audio and provided to the user.

[0556] Example 3: Creative idea proposal task

[0557] If a user types "Give me an idea for a science project," the generative AI model generates the idea "Build a small wind turbine using household materials," which is also provided to the user via voice synthesis.

[0558] In this way, each component works in cooperation with the others to efficiently and flexibly provide a variety of responses according to the text entered by the user.

[0559] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0560] Program processing flow and specific explanation

[0561] Step 1:

[0562] The user inputs text. For example, the user inputs "Translate 'Hello' to Japanese" into a text input field on the terminal. The input text is received by the terminal.

[0563] Step 2:

[0564] The device formats the received text into JSON format. Specifically, it uses JavaScript to convert the input text into an object format, and then uses the JSON.stringify function to convert it into a JSON string. This formatted data is sent as an HTTP POST request to the API endpoint.

[0565] Input: Raw text entered by the user ("Translate 'Hello' to Japanese")

[0566] Output: JSON formatted data

[0567] Step 3:

[0568] The server receives the request and parses the received JSON data. The server receives the request using the Python Flask framework, and parses the JSON data using request.get_json().

[0569] Input: JSON format data

[0570] Output: Parsed text ("Translate 'Hello' to Japanese")

[0571] Step 4:

[0572] The server analyzes the parsed text to determine whether it is a translation, summary, creative idea suggestion, or sentiment analysis, for example using a text analysis library (NLTK or spaCy).

[0573] Input: Parsed text

[0574] Output: The type of task (in this case, a translation task)

[0575] Step 5:

[0576] The server calls a generative AI model or emotion engine based on the identified task. For example, for a translation task, it sends a translation request to a generative AI model (e.g., GPT-3). In this case, it uses the GPT-3 API to send a request to translate "Hello" into Japanese.

[0577] Input: Task type (translation task) and source text ("Hello")

[0578] Output: Translation result("Hello")

[0579] Step 6:

[0580] The generated response text is then formatted into JSON again. The server then uses the jsonify function to convert the translation result into JSON format and sends it back as an HTTP response.

[0581] Input: Translation result("Hello")

[0582] Output: Response data in JSON format

[0583] Step 7:

[0584] The device analyzes the received response data and converts it into audio. The device uses JavaScript to parse the JSON data and converts the translation results into audio data using a speech synthesis API (such as Google Cloud Text-to-Speech). The generated audio data is played using the browser's Audio object.

[0585] Input: Response data in JSON format

[0586] Output: Audio data

[0587] Step 8:

[0588] The device plays the generated audio data through the speaker and provides it to the user as sound. Specifically, it plays the audio using the Audio.play() method.

[0589] Input: Audio data

[0590] Output: Speech provided to the user ("Hello")

[0591] The above is the specific processing flow of this system and the detailed operation of each step, which allows users to efficiently and flexibly obtain results for translation and other tasks.

[0592] (Application example 2)

[0593] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0594] Conventional food delivery services require users to input their menu items through unintuitive text input, and are unable to provide appropriate suggestions based on the user's emotions. This results in a poor user experience and reduces the appeal of the service.

[0595] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input text, means for receiving the text and transferring it to a generative AI model for analysis, means for the generative AI model to generate an appropriate response based on the text, means for converting the generated response into speech, means for providing the speech to the user, means for analyzing the user's emotions using the generative AI model and an emotion engine, and means for suggesting an appropriate menu to the user based on the emotion analysis. This allows the user to intuitively place an order using natural language and receive appropriate menu suggestions based on their emotions.

[0596] A "means for user text input" is a device or software that provides an interface for a user to input text in a natural language.

[0597] The "means for receiving said text and forwarding it to a generative AI model for analysis" refers to a function that sends user-entered text to a server and prepares the data to be analyzed by a generative AI model.

[0598] The "means by which the generative AI model generates an appropriate response based on the text" refers to the algorithm or process by which the generative AI model generates an appropriate reply or suggestion based on the text analyzed by the generative AI model.

[0599] "Means for converting the generated response into speech" refers to a technology or system that converts the text response generated by the generative AI model into speech data.

[0600] The "means for providing the audio to the user" refers to a device such as a speaker or earphone for playing back the data converted into audio to the user.

[0601] The "means for analyzing user emotions using the generative AI model and emotion engine" refers to a system that combines a generative AI model and an emotion analysis engine to identify emotions from the user's input text and adjust responses and suggestions based on the results.

[0602] The "means for suggesting an appropriate menu to a user based on the emotion analysis" is a system that automatically suggests an appropriate menu to a user in a food delivery service based on the user's emotions analyzed by an emotion analysis engine.

[0603] This invention is a system that processes user-entered text using a generative AI model and an emotion engine to support ordering and suggestions in food delivery services. This system consists of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0604] A user inputs text through their device. The device has a text input field, and when the user types, the text is received by the device. This text is then formatted into a standard data format such as JSON and sent to the server.

[0605] The server analyzes the data received from the device and determines whether it includes a translation task, a summary task, creative suggestions, or sentiment analysis. The generative AI model and the emotion engine work together to execute a specific task. Specifically, the server sends an analysis request to the generative AI model to request the necessary processing. The emotion engine also analyzes the user's sentiment from the input text, and the generative AI model adjusts its response accordingly.

[0606] The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. The emotion engine can recognize the user's emotions from a variety of input data, including text, voice, and facial expressions. For example, if the text entered by the user contains emotions such as "I had a tough day, recommend me something soothing," the emotion engine will identify emotions such as fatigue and anxiety, and the generative AI model will generate an appropriate response based on this.

[0607] The response text generated by the generative AI model is formatted in JSON and returned to the server. The server receives this generated response text and sends it back to the device. The device then uses speech synthesis technology to convert the received text into audio data. The text is converted into audio through a speech synthesis API, and the audio data is provided to the user through the speaker.

[0608] For example, if a user types "I want to order something spicy," the generative AI model will suggest "spicy dish suggestions," which will be received by the device and played aloud to the user via speech synthesis.

[0609] For example, if a user types, "I had a tough day, recommend me something soothing," the emotion engine will recognize emotions such as "fatigue" and "anxiety," and the generative AI model will generate an appropriate response, such as "recommended smoothies to relax you."

[0610] In this way, the system of the present invention allows users to efficiently obtain translations, summaries, and creative suggestions, as well as receive appropriate menu suggestions based on their current emotions, thereby improving the user experience in food delivery services and increasing their appeal.

[0611] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0612] Step 1:

[0613] The user inputs text into the terminal. The input text includes the user's order details or questions. For example, "I want to order something spicy." The input is formatted into a standard data format such as JSON. The input data is sent from the terminal to the server.

[0614] Step 2:

[0615] The server receives the text sent from the device, stores it in a database, and formats it appropriately for analysis. The server analyzes the text and determines whether it needs to be translated, summarized, suggested, or sentiment analyzed.

[0616] Step 3:

[0617] The server sends an analysis request to the generative AI model, which generates a highly accurate response based on the received text. For example, if a user types "I want to order something spicy," the generative AI model generates "spicy dish suggestions."

[0618] Step 4:

[0619] The server receives the response from the generative AI model and passes the text to the emotion engine to analyze the user's emotions. The emotion engine identifies the emotions contained in the text and adjusts the response accordingly. For example, for the input "I had a tough day, recommend me something soothing," the emotion engine identifies emotions such as "fatigue" and "anxiety."

[0620] Step 5:

[0621] The server combines the generated response text with the sentiment analysis results to generate optimal suggestions and responses, and the combined results are formatted in JSON format.

[0622] Step 6:

[0623] The server sends the formatted response text to the device, which then uses speech synthesis technology to provide a response to the user. The speech synthesis API converts the text into speech and provides it to the user as audio data through the speaker.

[0624] Step 7:

[0625] Users receive responses and suggestions through voice data, such as suggestions for spicy dishes or relaxing menu items, based on their emotions and the order they make.

[0626] In this way, a system is constructed that uses a generative AI model and emotion engine to perform advanced analysis and data calculations based on the user's input at each processing step, providing the user with an appropriate response.

[0627] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0628] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0629] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0630] [Third embodiment]

[0631] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0632] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0633] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0634] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0635] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0636] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0637] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0638] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0639] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0640] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0641] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0642] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0643] This invention relates to a communication support system that processes user-entered text using a generative AI model, translating, summarizing, and proposing creative ideas, and then providing the results as voice data. This system is comprised of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0644] The user inputs text through their device. The device has a text input field, and when the user inputs text, the text is received by the device. The received text is formatted into a standard data format such as JSON and then sent to the server.

[0645] The server analyzes the data received from the device and determines whether it corresponds to a translation task, a summarization task, or a creative idea suggestion. The generative AI model is then used for this processing. Specifically, the server sends an analysis request to the generative AI model, which then translates, summarizes, or generates creative ideas as needed.

[0646] The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. For example, when a user enters the text "Translate 'Hello' to Japanese," the generative AI model generates the translation result "Hello." This generated response text is then reformatted into JSON format and returned to the server.

[0647] The server receives the response text from the generative AI model and sends it to the device. The device uses speech synthesis technology to convert the received text into audio data. The text is converted into audio through a speech synthesis API, and the audio data is provided to the user through the speaker.

[0648] For example, if a user types "This is a long document that needs summarizing...", the generative AI model will summarize the long sentence and generate the summary text "This document needs summarizing". The device will receive this summary text and play it back to the user via speech synthesis.

[0649] If a user types "Give me an idea for a science project," the generative AI model will generate creative ideas, such as "Build a small wind turbine using household materials." The device will convert this idea into audio data and play it back to the user.

[0650] This allows users to efficiently meet a variety of needs, such as translation, summarization, and creative idea suggestions.The system of the present invention provides multiple functions in an integrated manner to support user communication, thereby resolving conventional problems.

[0651] The processing flow will be explained below.

[0652] Step 1:

[0653] A user enters text into an input field on a terminal. For example, "Translate 'Hello' to Japanese."

[0654] Step 2:

[0655] The device captures the user's input text and formats it into JSON format. For example, it generates data like {"action": "translate", "text": "Hello", "target_language": "Japanese"}.

[0656] Step 3:

[0657] The device sends the formatted data to the server, using an Internet connection to send a request to the server.

[0658] Step 4:

[0659] The server analyzes the data received from the device. It parses the received JSON data and checks that the "action" field is "translate".

[0660] Step 5:

[0661] The server transfers the analyzed data to the generative AI model. Specifically, it calls the generative AI model's API and requests a translation task. For example, make the following API call:

[0662] python

[0663] translation_result = ai_model.translate(text="Hello", target_language="Japanese")

[0664] Step 6:

[0665] The generative AI model translates the text "Hello" to "Hello." The translation result is returned to the server.

[0666] Step 7:

[0667] The server formats the translation results received from the AI ​​model into JSON format. For example, it generates data like {"translated_text": "Hello"}.

[0668] Step 8:

[0669] The server sends the formatted data to the device, again using the internet connection to return a response to the device.

[0670] Step 9:

[0671] The device parses the JSON data received from the server, interpreting the format {"translated_text": "Hello"} and extracting the contents of the "translated_text" field.

[0672] Step 10:

[0673] Convert text received by the device into speech. Use a speech synthesis API to convert text into speech data. Example: Use the following API:

[0674] python

[0675] synthesized_audio = text_to_speech("Hello")

[0676] Step 11:

[0677] The terminal plays the generated voice data to the user, and provides the voice to the user through the terminal's speaker.

[0678] Example 1

[0679] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0680] Conventional communication support systems are limited in the processing of text entered by users, making it difficult to quickly and accurately respond to diverse needs such as translation, summarization, and suggesting creative ideas. In addition, the series of processes required to provide generated responses as voice data are not integrated, leaving a need for an improved user experience.

[0681] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0682] In this invention, the server includes means for a user to input text, means for receiving the text and formatting it into a standard data format, means for transmitting the formatted data to the server, means for the server to analyze the received data and request processing from a generative AI model, means for the generative AI model to generate an appropriate response based on the text, means for receiving the generated response and transmitting it to a terminal, means for converting the received response into voice data, and means for providing the voice data to the user. This allows users to quickly and accurately respond to a variety of needs, such as translation, summarization, and creative idea suggestions.

[0683] A "user" is an individual or organization that inputs text and uses the system's functionality based on that input.

[0684] "Text" is a string of characters that a user inputs into the system.

[0685] A "terminal" is a device that allows a user to input text and processes it, such as receiving, sending, and formatting.

[0686] A "standard data format" is a common data format used to format text data, such as JSON.

[0687] A "server" is a computer system that receives data sent from a terminal, analyzes it, and requests processing from the generative AI model.

[0688] A "generative AI model" is an artificial intelligence model that translates text, summarizes it, suggests creative ideas, and more, based on pre-trained data.

[0689] "Means for requesting processing" refers to the method or technology by which the server sends an analysis request to the generated AI model.

[0690] A "response" is the textual result generated by a generative AI model.

[0691] "Audio data" is data obtained by converting text into audio.

[0692] "Speech synthesis" is a technology that converts text into audio data. For example, it uses a speech synthesis API.

[0693] "Means for providing" refers to a method or technology for delivering the generated voice data to the user.

[0694] This invention relates to a communication support system that processes user-entered text using a generative AI model, translating, summarizing, and proposing creative ideas, and then providing the results as voice data. This system is comprised of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0695] A user inputs text through their terminal. The terminal has a text input field, and when the user inputs text, the text is received by the terminal. For example, the user inputs "Translate 'Hello' to Japanese."

[0696] The terminal formats the received text into a standard data format such as JSON. For example, the input text is formatted into JSON as follows:

[0697] json

[0698] {

[0699] "text": "Translate 'Hello' to Japanese"

[0700] }

[0701] The formatted data is sent to a server, which analyzes the received data and determines whether it corresponds to a translation task, a summarization task, or a creative idea proposal task.

[0702] Once the analysis is complete, the server sends an analysis request to the generative AI model. The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. For example, in response to the prompt "Translate 'Hello' to Japanese," the generative AI model will generate the translation result "Hello." This generated response text is then reformatted into JSON format and returned to the server.

[0703] The server receives the response text from the generative AI model and sends it to the device. The device then uses speech synthesis technology to convert the received text into audio data. Specifically, the text is converted into audio through a speech synthesis API (e.g., Google Text-to-Speech API) and the audio data is provided to the user through a speaker. For example, if a user types "This is a long document that needs summarizing...", the generative AI model summarizes the long text and generates summary text that reads "This document needs summarizing." The device then receives this summary text and plays it back to the user through speech synthesis.

[0704] If a user types "Give me an idea for a science project," the generative AI model will generate creative ideas, such as "Build a small wind turbine using household materials." The device will convert this idea into audio data and play it back to the user.

[0705] This allows users to efficiently meet a variety of needs, such as translation, summarization, and creative idea suggestions.The system of the present invention provides multiple functions in an integrated manner and supports user communication, thereby resolving conventional problems.

[0706] The following are examples of prompt sentences:

[0707] 1. "Translate 'Hello' to Japanese" (translation task)

[0708] 2. "Summarize this document: This is a long document that needs summarizing..." (summarization task)

[0709] 3. "Give me an idea for a science project" (creative idea proposal task)

[0710] This allows users to quickly and accurately meet their various language and content needs.

[0711] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0712] Step 1:

[0713] The user enters text into the terminal

[0714] The user enters text into the device's text input field, for example, "Translate 'Hello' to Japanese." The entered text is stored in the device's memory.

[0715] Input: User-entered text "Translate 'Hello' to Japanese"

[0716] Output: Text "Translate 'Hello' to Japanese" stored in the device's memory

[0717] Step 2:

[0718] The device receives the text and formats it into JSON.

[0719] The terminal receives the input text and formats it into a standard data format (e.g., JSON), which ensures data consistency and streamlines subsequent processing.

[0720] Input: User-entered text "Translate 'Hello' to Japanese"

[0721] Output: JSON formatted data { "text": "Translate 'Hello' to Japanese"}

[0722] Step 3:

[0723] The terminal sends the formatted data to the server

[0724] The device sends data formatted in JSON to the server using a communication protocol such as an HTTP POST request.

[0725] Input: JSON format data { "text": "Translate 'Hello' to Japanese"}

[0726] Output: Data sent to the server { "text": "Translate 'Hello' to Japanese"}

[0727] Step 4:

[0728] The server analyzes the data and requests processing from the generative AI model.

[0729] The server analyzes the received data and determines that it is a translation task, then sends an analysis request to the generative AI model.

[0730] Input: Data received by the server { "text": "Translate 'Hello' to Japanese"}

[0731] Output: Parsing request sent to the generative AI model { "task": "translate", "text": "Hello", "target_language": "Japanese"}

[0732] Step 5:

[0733] A generative AI model generates response text

[0734] The generative AI model receives a request from the server and generates the Japanese translation of "Hello," which is "Konnichiwa." The generated response text is then reformatted into JSON and sent back to the server.

[0735] Input: Parsing request sent to the generative AI model { "task": "translate", "text": "Hello", "target_language": "Japanese"}

[0736] Output: Response text sent back to the server from the generative AI model { "translated_text": "Hello"}

[0737] Step 6:

[0738] The server receives the response text and sends it to the device.

[0739] The server sends the response text received from the generative AI model to the terminal.

[0740] Input: Response text returned by the generative AI model { "translated_text": "Hello"}

[0741] Output: Response text sent to the device { "translated_text": "Hello"}

[0742] Step 7:

[0743] The device converts the text into audio data and provides it to the user.

[0744] The device converts the received response text into voice data, and uses a speech synthesis API to convert "hello" into voice format and play it back to the user through the speaker.

[0745] Input: Response text received by the device { "translated_text": "Hello"}

[0746] Output: Voice data "Hello"

[0747] This allows users to efficiently receive text translations, summaries, and creative idea suggestions. For example, if a user types "This is a long document that needs summarizing...", the generative AI model will summarize the content and the device will provide the summary to the user via voice.

[0748] (Application example 1)

[0749] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0750] Conventional factory robots have difficulty understanding instructions and immediately providing specific work procedures and a summary of the situation when they receive them. Furthermore, as their use in multilingual environments increases, translation functions are often inadequate. This can lead to reduced work efficiency and communication problems.

[0751] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0752] In this invention, the server includes: a means for a user to input text; a means for receiving the text and transferring it to a generative AI model for analysis; a means for the generative AI model to generate an appropriate response based on the text; a means for converting the generated response into speech; and a means for providing the speech to the user. The system is installed in a factory robot, and the system generates a summary of work procedures and work status using the generative AI model in response to instructions from a worker and provides the summary by speech. This enables quick understanding of work instructions and status, and maintains high work efficiency even in a multilingual environment.

[0753] "User" refers to the worker who operates the factory robot and issues work instructions and requests information.

[0754] "Means for inputting text" refers to an interface that allows workers to input work instructions and information requests to the robot.

[0755] "Generative AI model" refers to an artificial intelligence model used to analyze input text and generate an appropriate response.

[0756] "Means for transmitting to the generative AI model for analysis" refers to the communications means for transmitting input text to the generative AI model for analysis.

[0757] "Means for generating an appropriate response" refers to the means by which a generative AI model analyzes text and translates, summarizes, or suggests work steps based on a specified task.

[0758] "Means for converting to speech" refers to speech synthesis technology used to convert text generated by a generative AI model into speech data.

[0759] "Means for providing audio to the user" refers to means for letting the worker hear the converted audio data through a speaker or the like.

[0760] The term "system installed on a factory robot" refers to a system including the above-mentioned means implemented on a factory robot, which enables interaction with workers.

[0761] "Work instructions" refer to instructions that specify the specific operations and tasks that workers will perform on factory robots.

[0762] "Summary of work procedures and status" refers to a concise explanation of the current work status and the next operation that a factory robot creates using a generative AI model based on instructions from a worker.

[0763] The system for implementing this invention is installed on a factory robot and utilizes a generative AI model in response to work instructions and information requests from workers. This system is configured as follows.

[0764] First, the user, a worker, inputs work instructions or information requests through a text input interface. The input text is received by the terminal, formatted into JSON format for analysis, and then sent to the server. The text input interface uses an HMI (Human Machine Interface) equipped with a touch panel and keyboard.

[0765] The server analyzes the received text and sends a request to the generative AI model based on its content. This generative AI model is capable of processing text in various languages ​​and generates highly accurate responses based on the Transformer architecture. Depending on the analysis, it translates, summarizes, or suggests work procedures. For example, if a prompt such as "Please tell me the steps required for the next process" is input, the generative AI model will generate the following response: "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."

[0766] The generated response is sent back to the server, which then sends it to the terminal. The terminal uses voice synthesis technology to convert the received text data into voice data. This voice synthesis technology uses gTTS (Google Text-to-Speech), and the generated voice data is provided to the worker through a speaker.

[0767] As a specific example, consider the case where a worker inputs the following text:

[0768] Please tell me the steps I need to take in the next step.

[0769] When this prompt is sent to the generative AI model, specific work procedures like those described above are generated and provided as voice data via the terminal.

[0770] This allows workers to receive the necessary information via voice at the appropriate time, enabling them to work efficiently.In addition, the system can translate and summarize with high accuracy in multilingual environments, enabling smooth communication between workers who speak different languages.

[0771] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0772] Step 1:

[0773] The user inputs text. The user uses a text input interface (such as a touch panel or keyboard) to input work instructions or information requests in text format into the terminal. For example, the user might input a prompt such as, "Please tell me the steps required for the next process."

[0774] Step 2:

[0775] The terminal receives the text and formats it into JSON format. The text entered by the user is received by the terminal and converted into a JSON format suitable for analysis. Specifically, the input text is processed into JSON data represented as key-value pairs. After this processing, the input text is formatted as "{"input": "Please tell me the steps required for the next step"}".

[0776] Step 3:

[0777] The device sends the formatted data to the server. The device then sends the generated JSON data to the server as an HTTP request, including appropriate header information (e.g., Content-Type: application / json).

[0778] Step 4:

[0779] The server analyzes the data and sends a request to the generative AI model. The server analyzes the received JSON data and determines the tasks required for the generative AI model. It then sends a request to the generative AI model API including the determined tasks (in this case, a proposed work procedure). Specifically, the server authenticates using an API key and sends a prompt to the generative AI model.

[0780] Step 5:

[0781] The generative AI model generates a response. The generative AI model analyzes the prompt received from the server and generates an appropriate response. For example, in response to the input, "Please tell me the steps required for the next process," it generates the response, "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."

[0782] Step 6:

[0783] The server receives the generated response and sends it to the terminal. After receiving the response text generated by the generative AI model, the server formats this data again into JSON format and sends it back to the terminal. The data to be sent is in the format "{"response": "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."}".

[0784] Step 7:

[0785] The device converts the response text into speech. The device passes the received response text to a speech synthesis API (e.g., gTTS) and converts it into audio data. The speech synthesis API analyzes the string and generates corresponding audio data (e.g., an mp3 file). The converted audio data is saved on the device.

[0786] Step 8:

[0787] The terminal provides the voice data to the user. The converted voice data is played back through the terminal's speaker, allowing the worker to receive a voice response. Specifically, the terminal plays back the generated voice data file and verbally tells the user the work procedure: "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."

[0788] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0789] This invention relates to a communication support system that processes user-entered text using a generative AI model and an emotion engine, translating, summarizing, suggesting creative ideas, and responding to the user's emotions, providing the results as voice data. This system is composed of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0790] A user inputs text through their device. The device has a text input field, and when the user inputs text, the text is received by the device. This text is then formatted into a standard data format such as JSON and sent to the server.

[0791] The server analyzes the data received from the device and determines whether it includes a translation task, a summary task, a creative idea suggestion, or sentiment analysis. The generative AI model and the emotion engine work together to execute a specific task. Specifically, the server sends an analysis request to the generative AI model to request the necessary processing. The emotion engine also analyzes the user's sentiment from the input text, and the generative AI model adjusts its response accordingly.

[0792] The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. The emotion engine can recognize the user's emotions from a variety of input data, including text, voice, and facial expressions. For example, if the text entered by the user contains the emotion "I'm feeling sad today," the emotion engine will identify the emotion "sad," and the generative AI model will generate an appropriate response based on this.

[0793] The response text generated by the generative AI model is formatted in JSON and returned to the server. The server receives this generated response text and sends it back to the device. The device then uses speech synthesis technology to convert the received text into audio data. The text is converted into audio through a speech synthesis API, and the audio data is provided to the user through the speaker.

[0794] For example, if a user types "Translate 'Hello' to Japanese," the generative AI model will generate the translation "Hello." This translation result is received by the device and played back to the user via speech synthesis.

[0795] Also, if a user types, "I had a bad day today," the emotion engine recognizes the emotion "negative," and the generative AI model generates an appropriate response, such as a message like, "I'm sorry to hear that. Would you like to talk about it?"

[0796] Furthermore, if a user types "Give me an idea for a science project," the generative AI model will generate the creative idea "Build a small wind turbine using household materials," which is also provided to the user via voice synthesis.

[0797] In this way, the system of the present invention allows users to efficiently translate, summarize, suggest creative ideas, and even receive emotional responses, thereby solving the problems of the past.

[0798] The processing flow will be explained below.

[0799] Step 1:

[0800] A user enters text into an input field on a terminal, for example, "I had a bad day today."

[0801] Step 2:

[0802] The device captures the user's input text and formats it into JSON, e.g., {"text": "I had a bad day today"}.

[0803] Step 3:

[0804] The device sends the formatted data to the server, using an Internet connection to send a request to the server.

[0805] Step 4:

[0806] The server analyzes the data received from the device. It parses the received JSON data and extracts the text portion.

[0807] Step 5:

[0808] The server transfers the analyzed data to the emotion engine. Specifically, it calls the emotion engine's API and requests emotion analysis. For example, make the following API call:

[0809] python

[0810] emotion_result = emotion_engine.analyze(text="I had a bad day today")

[0811] Step 6:

[0812] The emotion engine analyzes the text "I had a bad day today" and recognizes the emotion "negative." This emotion information is returned to the server.

[0813] Step 7:

[0814] The server transfers the emotion information received from the emotion engine to the generative AI model. At the same time, it also sends the original text. For example, make the following API call:

[0815] python

[0816] response = ai_model.generate_response(text="I had a bad day today", emotion="negative")

[0817] Step 8:

[0818] A generative AI model generates an appropriate response based on the text and sentiment information, e.g., "I'm sorry to hear that. Would you like to talk about it?"

[0819] Step 9:

[0820] The server formats the response received from the generated AI model into JSON format. For example, it generates data like {"response_text": "I'm sorry to hear that. Would you like to talk about it?"}.

[0821] Step 10:

[0822] The server then sends the formatted data to the device, again using the internet connection to return a response to the device.

[0823] Step 11:

[0824] The device parses the JSON data received from the server, interpreting the format {"response_text": "I'm sorry to hear that. Would you like to talk about it?"} and extracts the text portion.

[0825] Step 12:

[0826] Convert text received by the device into speech. Use a speech synthesis API to convert text into speech data. Example: Use the following API:

[0827] python

[0828] synthesized_audio = text_to_speech("I'm sorry to hear that. Would you like to talk about it?")

[0829] Step 13:

[0830] The terminal plays the generated voice data to the user, and provides the voice to the user through the terminal's speaker.

[0831] This allows the user to receive an appropriate voice response that corresponds to their own emotions.

[0832] Example 2

[0833] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0834] Conventional communication support systems have a limited ability to properly analyze text entered by users and provide responses based on that content. Furthermore, they are unable to generate responses that reflect the user's emotions, making it difficult to improve the quality of communication. Furthermore, functions such as text translation and summarization are often provided separately, making them less convenient as an integrated system.

[0835] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for a user to input text, means for receiving the text, shaping it into a data format, and then transmitting it to the server for analysis, means for the server to analyze the text and determine a task, means for the generative AI model and emotion engine to perform appropriate processing based on the task, means for converting the generated response into speech, and means for providing the speech to the user. This makes it possible to automatically provide appropriate translations, summaries, suggestions for creative ideas, and even responses according to emotions based on the text entered by the user.

[0836] "User" refers to any individual or corporation that uses this system.

[0837] "Means for inputting text" refers to an interface for a user to input text information, such as a keyboard or a touch screen.

[0838] "Data format" refers to a format for formatting information based on certain rules, and includes, for example, JSON format.

[0839] A "server" refers to a computer system that provides services to other computers over a network.

[0840] A "generative AI model" refers to an artificial intelligence model that is trained using large amounts of text data and generates and analyzes text.

[0841] "Emotion engine" refers to a system for analyzing a user's emotions from input text.

[0842] "Means for converting a response into voice" refers to technology for converting text data into voice data, including, for example, a voice synthesis API.

[0843] A "task" refers to a specific process or function that the system must perform, such as translation, summarization, creative idea suggestion, or sentiment analysis.

[0844] "Analysis" refers to the process of analyzing input text data and clarifying its content and intent.

[0845] This invention relates to a communication support system that performs advanced processing of text entered by a user, and provides translation, summarization, creative idea suggestions, and even responses tailored to the user's emotions. This system consists of three main components: the user, the terminal, and the server. Each component and its operation are described in detail below.

[0846] User operations and device roles

[0847] A user inputs text through a terminal such as a personal computer or smartphone. The terminal is provided with a text input field, and when the user inputs text, the text is received by the terminal. The input text is formatted into a standard data format such as JSON. This formatted data is then sent to the server via an HTTP request.

[0848] Data analysis and model execution on the server

[0849] The server analyzes the received data and determines whether it corresponds to a translation task, a summarization task, a creative idea suggestion, or sentiment analysis. This determination process uses a text analysis library (e.g., NLTK or spaCy). After determining the specific task, the server invokes a generative AI model and sentiment engine to perform the necessary processing.

[0850] Generative AI models are pre-trained using large amounts of text data, enabling accurate analysis and response generation. For example, OpenAI's GPT-3 is used. Meanwhile, emotion engines are systems that recognize user emotions from input text, such as IBM Watson's Tone Analyzer.

[0851] Response generation and information return

[0852] The response text generated by the generative AI model is again formatted in JSON format. This response data is sent from the server to the device. The device then uses speech synthesis technology to convert the received response text into audio data. For example, Google Cloud Text-to-Speech API is used. The converted audio data is then provided to the user via the speaker.

[0853] Specific examples

[0854] Below is a concrete example of a prompt entered by a user and the system's response to it.

[0855] Example 1: Translation task

[0856] When a user types "Translate 'Hello' to Japanese," the device sends the text to the server, and the generative AI model generates the translation result "Hello," which is then converted into speech on the device and played back to the user.

[0857] Example 2: Emotional response task

[0858] If a user types "I had a bad day today," the emotion engine recognizes the emotion "negative," and the generative AI model generates the corresponding message, "I'm sorry to hear that. Would you like to talk about it?" This message is also converted to audio and provided to the user.

[0859] Example 3: Creative idea proposal task

[0860] If a user types "Give me an idea for a science project," the generative AI model generates the idea "Build a small wind turbine using household materials," which is also provided to the user via voice synthesis.

[0861] In this way, each component works in cooperation with the others to efficiently and flexibly provide a variety of responses according to the text entered by the user.

[0862] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0863] Program processing flow and specific explanation

[0864] Step 1:

[0865] The user inputs text. For example, the user inputs "Translate 'Hello' to Japanese" into a text input field on the terminal. The input text is received by the terminal.

[0866] Step 2:

[0867] The device formats the received text into JSON format. Specifically, it uses JavaScript to convert the input text into an object format, and then uses the JSON.stringify function to convert it into a JSON string. This formatted data is sent as an HTTP POST request to the API endpoint.

[0868] Input: Raw text entered by the user ("Translate 'Hello' to Japanese")

[0869] Output: JSON formatted data

[0870] Step 3:

[0871] The server receives the request and parses the received JSON data. The server receives the request using the Python Flask framework, and parses the JSON data using request.get_json().

[0872] Input: JSON format data

[0873] Output: Parsed text ("Translate 'Hello' to Japanese")

[0874] Step 4:

[0875] The server analyzes the parsed text to determine whether it is a translation, summary, creative idea suggestion, or sentiment analysis, for example using a text analysis library (NLTK or spaCy).

[0876] Input: Parsed text

[0877] Output: The type of task (in this case, a translation task)

[0878] Step 5:

[0879] The server calls a generative AI model or emotion engine based on the identified task. For example, for a translation task, it sends a translation request to a generative AI model (e.g., GPT-3). In this case, it uses the GPT-3 API to send a request to translate "Hello" into Japanese.

[0880] Input: Task type (translation task) and source text ("Hello")

[0881] Output: Translation result("Hello")

[0882] Step 6:

[0883] The generated response text is then formatted into JSON again. The server then uses the jsonify function to convert the translation result into JSON format and sends it back as an HTTP response.

[0884] Input: Translation result("Hello")

[0885] Output: Response data in JSON format

[0886] Step 7:

[0887] The device analyzes the received response data and converts it into audio. The device uses JavaScript to parse the JSON data and converts the translation results into audio data using a speech synthesis API (such as Google Cloud Text-to-Speech). The generated audio data is played using the browser's Audio object.

[0888] Input: Response data in JSON format

[0889] Output: Audio data

[0890] Step 8:

[0891] The device plays the generated audio data through the speaker and provides it to the user as sound. Specifically, it plays the audio using the Audio.play() method.

[0892] Input: Audio data

[0893] Output: Speech provided to the user ("Hello")

[0894] The above is the specific processing flow of this system and the detailed operation of each step, which allows users to efficiently and flexibly obtain results for translation and other tasks.

[0895] (Application example 2)

[0896] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0897] Conventional food delivery services require users to input their menu items through unintuitive text input, and are unable to provide appropriate suggestions based on the user's emotions. This results in a poor user experience and reduces the appeal of the service.

[0898] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input text, means for receiving the text and transferring it to a generative AI model for analysis, means for the generative AI model to generate an appropriate response based on the text, means for converting the generated response into speech, means for providing the speech to the user, means for analyzing the user's emotions using the generative AI model and an emotion engine, and means for suggesting an appropriate menu to the user based on the emotion analysis. This allows the user to intuitively place an order using natural language and receive appropriate menu suggestions based on their emotions.

[0899] A "means for user text input" is a device or software that provides an interface for a user to input text in a natural language.

[0900] The "means for receiving said text and forwarding it to a generative AI model for analysis" refers to a function that sends user-entered text to a server and prepares the data to be analyzed by a generative AI model.

[0901] The "means by which the generative AI model generates an appropriate response based on the text" refers to the algorithm or process by which the generative AI model generates an appropriate reply or suggestion based on the text analyzed by the generative AI model.

[0902] "Means for converting the generated response into speech" refers to a technology or system that converts the text response generated by the generative AI model into speech data.

[0903] The "means for providing the audio to the user" refers to a device such as a speaker or earphone for playing back the data converted into audio to the user.

[0904] The "means for analyzing user emotions using the generative AI model and emotion engine" refers to a system that combines a generative AI model and an emotion analysis engine to identify emotions from the user's input text and adjust responses and suggestions based on the results.

[0905] The "means for suggesting an appropriate menu to a user based on the emotion analysis" is a system that automatically suggests an appropriate menu to a user in a food delivery service based on the user's emotions analyzed by an emotion analysis engine.

[0906] This invention is a system that processes user-entered text using a generative AI model and an emotion engine to support ordering and suggestions in food delivery services. This system consists of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0907] A user inputs text through their device. The device has a text input field, and when the user types, the text is received by the device. This text is then formatted into a standard data format such as JSON and sent to the server.

[0908] The server analyzes the data received from the device and determines whether it includes a translation task, a summary task, creative suggestions, or sentiment analysis. The generative AI model and the emotion engine work together to execute a specific task. Specifically, the server sends an analysis request to the generative AI model to request the necessary processing. The emotion engine also analyzes the user's sentiment from the input text, and the generative AI model adjusts its response accordingly.

[0909] The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. The emotion engine can recognize the user's emotions from a variety of input data, including text, voice, and facial expressions. For example, if the text entered by the user contains emotions such as "I had a tough day, recommend me something soothing," the emotion engine will identify emotions such as fatigue and anxiety, and the generative AI model will generate an appropriate response based on this.

[0910] The response text generated by the generative AI model is formatted in JSON and returned to the server. The server receives this generated response text and sends it back to the device. The device then uses speech synthesis technology to convert the received text into audio data. The text is converted into audio through a speech synthesis API, and the audio data is provided to the user through the speaker.

[0911] For example, if a user types "I want to order something spicy," the generative AI model will suggest "spicy dish suggestions," which will be received by the device and played aloud to the user via speech synthesis.

[0912] For example, if a user types, "I had a tough day, recommend me something soothing," the emotion engine will recognize emotions such as "fatigue" and "anxiety," and the generative AI model will generate an appropriate response, such as "recommended smoothies to relax you."

[0913] In this way, the system of the present invention allows users to efficiently obtain translations, summaries, and creative suggestions, as well as receive appropriate menu suggestions based on their current emotions, thereby improving the user experience in food delivery services and increasing their appeal.

[0914] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0915] Step 1:

[0916] The user inputs text into the terminal. The input text includes the user's order details or questions. For example, "I want to order something spicy." The input is formatted into a standard data format such as JSON. The input data is sent from the terminal to the server.

[0917] Step 2:

[0918] The server receives the text sent from the device, stores it in a database, and formats it appropriately for analysis. The server analyzes the text and determines whether it needs to be translated, summarized, suggested, or sentiment analyzed.

[0919] Step 3:

[0920] The server sends an analysis request to the generative AI model, which generates a highly accurate response based on the received text. For example, if a user types "I want to order something spicy," the generative AI model generates "spicy dish suggestions."

[0921] Step 4:

[0922] The server receives the response from the generative AI model and passes the text to the emotion engine to analyze the user's emotions. The emotion engine identifies the emotions contained in the text and adjusts the response accordingly. For example, for the input "I had a tough day, recommend me something soothing," the emotion engine identifies emotions such as "fatigue" and "anxiety."

[0923] Step 5:

[0924] The server combines the generated response text with the sentiment analysis results to generate optimal suggestions and responses, and the combined results are formatted in JSON format.

[0925] Step 6:

[0926] The server sends the formatted response text to the device, which then uses speech synthesis technology to provide a response to the user. The speech synthesis API converts the text into speech and provides it to the user as audio data through the speaker.

[0927] Step 7:

[0928] Users receive responses and suggestions through voice data, such as suggestions for spicy dishes or relaxing menu items, based on their emotions and the order they make.

[0929] In this way, a system is constructed that uses a generative AI model and emotion engine to perform advanced analysis and data calculations based on the user's input at each processing step, providing the user with an appropriate response.

[0930] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0931] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0932] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0933] [Fourth embodiment]

[0934] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0935] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0936] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0937] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0938] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0939] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0940] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0941] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0942] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0943] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0944] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0945] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0946] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0947] This invention relates to a communication support system that processes user-entered text using a generative AI model, translating, summarizing, and proposing creative ideas, and then providing the results as voice data. This system is comprised of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0948] The user inputs text through their device. The device has a text input field, and when the user inputs text, the text is received by the device. The received text is formatted into a standard data format such as JSON and then sent to the server.

[0949] The server analyzes the data received from the device and determines whether it corresponds to a translation task, a summarization task, or a creative idea suggestion. The generative AI model is then used for this processing. Specifically, the server sends an analysis request to the generative AI model, which then translates, summarizes, or generates creative ideas as needed.

[0950] The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. For example, when a user enters the text "Translate 'Hello' to Japanese," the generative AI model generates the translation result "Hello." This generated response text is then reformatted into JSON format and returned to the server.

[0951] The server receives the response text from the generative AI model and sends it to the device. The device uses speech synthesis technology to convert the received text into audio data. The text is converted into audio through a speech synthesis API, and the audio data is provided to the user through the speaker.

[0952] For example, if a user types "This is a long document that needs summarizing...", the generative AI model will summarize the long sentence and generate the summary text "This document needs summarizing". The device will receive this summary text and play it back to the user via speech synthesis.

[0953] If a user types "Give me an idea for a science project," the generative AI model will generate creative ideas, such as "Build a small wind turbine using household materials." The device will convert this idea into audio data and play it back to the user.

[0954] This allows users to efficiently meet a variety of needs, such as translation, summarization, and creative idea suggestions.The system of the present invention provides multiple functions in an integrated manner to support user communication, thereby resolving conventional problems.

[0955] The processing flow will be explained below.

[0956] Step 1:

[0957] A user enters text into an input field on a terminal. For example, "Translate 'Hello' to Japanese."

[0958] Step 2:

[0959] The device captures the user's input text and formats it into JSON format. For example, it generates data like {"action": "translate", "text": "Hello", "target_language": "Japanese"}.

[0960] Step 3:

[0961] The device sends the formatted data to the server, using an Internet connection to send a request to the server.

[0962] Step 4:

[0963] The server analyzes the data received from the device. It parses the received JSON data and checks that the "action" field is "translate".

[0964] Step 5:

[0965] The server transfers the analyzed data to the generative AI model. Specifically, it calls the generative AI model's API and requests a translation task. For example, make the following API call:

[0966] python

[0967] translation_result = ai_model.translate(text="Hello", target_language="Japanese")

[0968] Step 6:

[0969] The generative AI model translates the text "Hello" to "Hello." The translation result is returned to the server.

[0970] Step 7:

[0971] The server formats the translation results received from the AI ​​model into JSON format. For example, it generates data like {"translated_text": "Hello"}.

[0972] Step 8:

[0973] The server sends the formatted data to the device, again using the internet connection to return a response to the device.

[0974] Step 9:

[0975] The device parses the JSON data received from the server, interpreting the format {"translated_text": "Hello"} and extracting the contents of the "translated_text" field.

[0976] Step 10:

[0977] Convert text received by the device into speech. Use a speech synthesis API to convert text into speech data. Example: Use the following API:

[0978] python

[0979] synthesized_audio = text_to_speech("Hello")

[0980] Step 11:

[0981] The terminal plays the generated voice data to the user, and provides the voice to the user through the terminal's speaker.

[0982] Example 1

[0983] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0984] Conventional communication support systems are limited in the processing of text entered by users, making it difficult to quickly and accurately respond to diverse needs such as translation, summarization, and suggesting creative ideas. In addition, the series of processes required to provide generated responses as voice data are not integrated, leaving a need for an improved user experience.

[0985] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0986] In this invention, the server includes means for a user to input text, means for receiving the text and formatting it into a standard data format, means for transmitting the formatted data to the server, means for the server to analyze the received data and request processing from a generative AI model, means for the generative AI model to generate an appropriate response based on the text, means for receiving the generated response and transmitting it to a terminal, means for converting the received response into voice data, and means for providing the voice data to the user. This allows users to quickly and accurately respond to a variety of needs, such as translation, summarization, and creative idea suggestions.

[0987] A "user" is an individual or organization that inputs text and uses the system's functionality based on that input.

[0988] "Text" is a string of characters that a user inputs into the system.

[0989] A "terminal" is a device that allows a user to input text and processes it, such as receiving, sending, and formatting.

[0990] A "standard data format" is a common data format used to format text data, such as JSON.

[0991] A "server" is a computer system that receives data sent from a terminal, analyzes it, and requests processing from the generative AI model.

[0992] A "generative AI model" is an artificial intelligence model that translates text, summarizes it, suggests creative ideas, and more, based on pre-trained data.

[0993] "Means for requesting processing" refers to the method or technology by which the server sends an analysis request to the generated AI model.

[0994] A "response" is the textual result generated by a generative AI model.

[0995] "Audio data" is data obtained by converting text into audio.

[0996] "Speech synthesis" is a technology that converts text into audio data. For example, it uses a speech synthesis API.

[0997] "Means for providing" refers to a method or technology for delivering the generated voice data to the user.

[0998] This invention relates to a communication support system that processes user-entered text using a generative AI model, translating, summarizing, and proposing creative ideas, and then providing the results as voice data. This system is comprised of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[0999] A user inputs text through their terminal. The terminal has a text input field, and when the user inputs text, the text is received by the terminal. For example, the user inputs "Translate 'Hello' to Japanese."

[1000] The terminal formats the received text into a standard data format such as JSON. For example, the input text is formatted into JSON as follows:

[1001] json

[1002] {

[1003] "text": "Translate 'Hello' to Japanese"

[1004] }

[1005] The formatted data is sent to a server, which analyzes the received data and determines whether it corresponds to a translation task, a summarization task, or a creative idea proposal task.

[1006] Once the analysis is complete, the server sends an analysis request to the generative AI model. The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. For example, in response to the prompt "Translate 'Hello' to Japanese," the generative AI model will generate the translation result "Hello." This generated response text is then reformatted into JSON format and returned to the server.

[1007] The server receives the response text from the generative AI model and sends it to the device. The device then uses speech synthesis technology to convert the received text into audio data. Specifically, the text is converted into audio through a speech synthesis API (e.g., Google Text-to-Speech API) and the audio data is provided to the user through a speaker. For example, if a user types "This is a long document that needs summarizing...", the generative AI model summarizes the long text and generates summary text that reads "This document needs summarizing." The device then receives this summary text and plays it back to the user through speech synthesis.

[1008] If a user types "Give me an idea for a science project," the generative AI model will generate creative ideas, such as "Build a small wind turbine using household materials." The device will convert this idea into audio data and play it back to the user.

[1009] This allows users to efficiently meet a variety of needs, such as translation, summarization, and creative idea suggestions.The system of the present invention provides multiple functions in an integrated manner and supports user communication, thereby resolving conventional problems.

[1010] The following are examples of prompt sentences:

[1011] 1. "Translate 'Hello' to Japanese" (translation task)

[1012] 2. "Summarize this document: This is a long document that needs summarizing..." (summarization task)

[1013] 3. "Give me an idea for a science project" (creative idea proposal task)

[1014] This allows users to quickly and accurately meet their various language and content needs.

[1015] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1016] Step 1:

[1017] The user enters text into the terminal

[1018] The user enters text into the device's text input field, for example, "Translate 'Hello' to Japanese." The entered text is stored in the device's memory.

[1019] Input: User-entered text "Translate 'Hello' to Japanese"

[1020] Output: Text "Translate 'Hello' to Japanese" stored in the device's memory

[1021] Step 2:

[1022] The device receives the text and formats it into JSON.

[1023] The terminal receives the input text and formats it into a standard data format (e.g., JSON), which ensures data consistency and streamlines subsequent processing.

[1024] Input: User-entered text "Translate 'Hello' to Japanese"

[1025] Output: JSON formatted data { "text": "Translate 'Hello' to Japanese"}

[1026] Step 3:

[1027] The terminal sends the formatted data to the server

[1028] The device sends data formatted in JSON to the server using a communication protocol such as an HTTP POST request.

[1029] Input: JSON format data { "text": "Translate 'Hello' to Japanese"}

[1030] Output: Data sent to the server { "text": "Translate 'Hello' to Japanese"}

[1031] Step 4:

[1032] The server analyzes the data and requests processing from the generative AI model.

[1033] The server analyzes the received data and determines that it is a translation task, then sends an analysis request to the generative AI model.

[1034] Input: Data received by the server { "text": "Translate 'Hello' to Japanese"}

[1035] Output: Parsing request sent to the generative AI model { "task": "translate", "text": "Hello", "target_language": "Japanese"}

[1036] Step 5:

[1037] A generative AI model generates response text

[1038] The generative AI model receives a request from the server and generates the Japanese translation of "Hello," which is "Konnichiwa." The generated response text is then reformatted into JSON and sent back to the server.

[1039] Input: Parsing request sent to the generative AI model { "task": "translate", "text": "Hello", "target_language": "Japanese"}

[1040] Output: Response text sent back to the server from the generative AI model { "translated_text": "Hello"}

[1041] Step 6:

[1042] The server receives the response text and sends it to the device.

[1043] The server sends the response text received from the generative AI model to the terminal.

[1044] Input: Response text returned by the generative AI model { "translated_text": "Hello"}

[1045] Output: Response text sent to the device { "translated_text": "Hello"}

[1046] Step 7:

[1047] The device converts the text into audio data and provides it to the user.

[1048] The device converts the received response text into voice data, and uses a speech synthesis API to convert "hello" into voice format and play it back to the user through the speaker.

[1049] Input: Response text received by the device { "translated_text": "Hello"}

[1050] Output: Voice data "Hello"

[1051] This allows users to efficiently receive text translations, summaries, and creative idea suggestions. For example, if a user types "This is a long document that needs summarizing...", the generative AI model will summarize the content and the device will provide the summary to the user via voice.

[1052] (Application example 1)

[1053] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1054] Conventional factory robots have difficulty understanding instructions and immediately providing specific work procedures and a summary of the situation when they receive them. Furthermore, as their use in multilingual environments increases, translation functions are often inadequate. This can lead to reduced work efficiency and communication problems.

[1055] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1056] In this invention, the server includes: a means for a user to input text; a means for receiving the text and transferring it to a generative AI model for analysis; a means for the generative AI model to generate an appropriate response based on the text; a means for converting the generated response into speech; and a means for providing the speech to the user. The system is installed in a factory robot, and the system generates a summary of work procedures and work status using the generative AI model in response to instructions from a worker and provides the summary by speech. This enables quick understanding of work instructions and status, and maintains high work efficiency even in a multilingual environment.

[1057] "User" refers to the worker who operates the factory robot and issues work instructions and requests information.

[1058] "Means for inputting text" refers to an interface that allows workers to input work instructions and information requests to the robot.

[1059] "Generative AI model" refers to an artificial intelligence model used to analyze input text and generate an appropriate response.

[1060] "Means for transmitting to the generative AI model for analysis" refers to the communications means for transmitting input text to the generative AI model for analysis.

[1061] "Means for generating an appropriate response" refers to the means by which a generative AI model analyzes text and translates, summarizes, or suggests work steps based on a specified task.

[1062] "Means for converting to speech" refers to speech synthesis technology used to convert text generated by a generative AI model into speech data.

[1063] "Means for providing audio to the user" refers to means for letting the worker hear the converted audio data through a speaker or the like.

[1064] The term "system installed on a factory robot" refers to a system including the above-mentioned means implemented on a factory robot, which enables interaction with workers.

[1065] "Work instructions" refer to instructions that specify the specific operations and tasks that workers will perform on factory robots.

[1066] "Summary of work procedures and status" refers to a concise explanation of the current work status and the next operation that a factory robot creates using a generative AI model based on instructions from a worker.

[1067] The system for implementing this invention is installed on a factory robot and utilizes a generative AI model in response to work instructions and information requests from workers. This system is configured as follows.

[1068] First, the user, a worker, inputs work instructions or information requests through a text input interface. The input text is received by the terminal, formatted into JSON format for analysis, and then sent to the server. The text input interface uses an HMI (Human Machine Interface) equipped with a touch panel and keyboard.

[1069] The server analyzes the received text and sends a request to the generative AI model based on its content. This generative AI model is capable of processing text in various languages ​​and generates highly accurate responses based on the Transformer architecture. Depending on the analysis, it translates, summarizes, or suggests work procedures. For example, if a prompt such as "Please tell me the steps required for the next process" is input, the generative AI model will generate the following response: "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."

[1070] The generated response is sent back to the server, which then sends it to the terminal. The terminal uses voice synthesis technology to convert the received text data into voice data. This voice synthesis technology uses gTTS (Google Text-to-Speech), and the generated voice data is provided to the worker through a speaker.

[1071] As a specific example, consider the case where a worker inputs the following text:

[1072] Please tell me the steps I need to take in the next step.

[1073] When this prompt is sent to the generative AI model, specific work procedures like those described above are generated and provided as voice data via the terminal.

[1074] This allows workers to receive the necessary information via voice at the appropriate time, enabling them to work efficiently.In addition, the system can translate and summarize with high accuracy in multilingual environments, enabling smooth communication between workers who speak different languages.

[1075] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1076] Step 1:

[1077] The user inputs text. The user uses a text input interface (such as a touch panel or keyboard) to input work instructions or information requests in text format into the terminal. For example, the user might input a prompt such as, "Please tell me the steps required for the next process."

[1078] Step 2:

[1079] The terminal receives the text and formats it into JSON format. The text entered by the user is received by the terminal and converted into a JSON format suitable for analysis. Specifically, the input text is processed into JSON data represented as key-value pairs. After this processing, the input text is formatted as "{"input": "Please tell me the steps required for the next step"}".

[1080] Step 3:

[1081] The device sends the formatted data to the server. The device then sends the generated JSON data to the server as an HTTP request, including appropriate header information (e.g., Content-Type: application / json).

[1082] Step 4:

[1083] The server analyzes the data and sends a request to the generative AI model. The server analyzes the received JSON data and determines the tasks required for the generative AI model. It then sends a request to the generative AI model API including the determined tasks (in this case, a proposed work procedure). Specifically, the server authenticates using an API key and sends a prompt to the generative AI model.

[1084] Step 5:

[1085] The generative AI model generates a response. The generative AI model analyzes the prompt received from the server and generates an appropriate response. For example, in response to the input, "Please tell me the steps required for the next process," it generates the response, "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."

[1086] Step 6:

[1087] The server receives the generated response and sends it to the terminal. After receiving the response text generated by the generative AI model, the server formats this data again into JSON format and sends it back to the terminal. The data to be sent is in the format "{"response": "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."}".

[1088] Step 7:

[1089] The device converts the response text into speech. The device passes the received response text to a speech synthesis API (e.g., gTTS) and converts it into audio data. The speech synthesis API analyzes the string and generates corresponding audio data (e.g., an mp3 file). The converted audio data is saved on the device.

[1090] Step 8:

[1091] The terminal provides the voice data to the user. The converted voice data is played back through the terminal's speaker, allowing the worker to receive a voice response. Specifically, the terminal plays back the generated voice data file and verbally tells the user the work procedure: "1. Place the material on the supply line. 2. Start the machine and make any necessary adjustments. 3. Inspect the sample."

[1092] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1093] This invention relates to a communication support system that processes user-entered text using a generative AI model and an emotion engine, translating, summarizing, suggesting creative ideas, and responding to the user's emotions, providing the results as voice data. This system is composed of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[1094] A user inputs text through their device. The device has a text input field, and when the user inputs text, the text is received by the device. This text is then formatted into a standard data format such as JSON and sent to the server.

[1095] The server analyzes the data received from the device and determines whether it includes a translation task, a summary task, a creative idea suggestion, or sentiment analysis. The generative AI model and the emotion engine work together to execute a specific task. Specifically, the server sends an analysis request to the generative AI model to request the necessary processing. The emotion engine also analyzes the user's sentiment from the input text, and the generative AI model adjusts its response accordingly.

[1096] The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. The emotion engine can recognize the user's emotions from a variety of input data, including text, voice, and facial expressions. For example, if the text entered by the user contains the emotion "I'm feeling sad today," the emotion engine will identify the emotion "sad," and the generative AI model will generate an appropriate response based on this.

[1097] The response text generated by the generative AI model is formatted in JSON and returned to the server. The server receives this generated response text and sends it back to the device. The device then uses speech synthesis technology to convert the received text into audio data. The text is converted into audio through a speech synthesis API, and the audio data is provided to the user through the speaker.

[1098] For example, if a user types "Translate 'Hello' to Japanese," the generative AI model will generate the translation "Hello." This translation result is received by the device and played back to the user via speech synthesis.

[1099] Also, if a user types, "I had a bad day today," the emotion engine recognizes the emotion "negative," and the generative AI model generates an appropriate response, such as a message like, "I'm sorry to hear that. Would you like to talk about it?"

[1100] Furthermore, if a user types "Give me an idea for a science project," the generative AI model will generate the creative idea "Build a small wind turbine using household materials," which is also provided to the user via voice synthesis.

[1101] In this way, the system of the present invention allows users to efficiently translate, summarize, suggest creative ideas, and even receive emotional responses, thereby solving the problems of the past.

[1102] The processing flow will be explained below.

[1103] Step 1:

[1104] A user enters text into an input field on a terminal, for example, "I had a bad day today."

[1105] Step 2:

[1106] The device captures the user's input text and formats it into JSON, e.g., {"text": "I had a bad day today"}.

[1107] Step 3:

[1108] The device sends the formatted data to the server, using an Internet connection to send a request to the server.

[1109] Step 4:

[1110] The server analyzes the data received from the device. It parses the received JSON data and extracts the text portion.

[1111] Step 5:

[1112] The server transfers the analyzed data to the emotion engine. Specifically, it calls the emotion engine's API and requests emotion analysis. For example, make the following API call:

[1113] python

[1114] emotion_result = emotion_engine.analyze(text="I had a bad day today")

[1115] Step 6:

[1116] The emotion engine analyzes the text "I had a bad day today" and recognizes the emotion "negative." This emotion information is returned to the server.

[1117] Step 7:

[1118] The server transfers the emotion information received from the emotion engine to the generative AI model. At the same time, it also sends the original text. For example, make the following API call:

[1119] python

[1120] response = ai_model.generate_response(text="I had a bad day today", emotion="negative")

[1121] Step 8:

[1122] A generative AI model generates an appropriate response based on the text and sentiment information, e.g., "I'm sorry to hear that. Would you like to talk about it?"

[1123] Step 9:

[1124] The server formats the response received from the generated AI model into JSON format. For example, it generates data like {"response_text": "I'm sorry to hear that. Would you like to talk about it?"}.

[1125] Step 10:

[1126] The server then sends the formatted data to the device, again using the internet connection to return a response to the device.

[1127] Step 11:

[1128] The device parses the JSON data received from the server, interpreting the format {"response_text": "I'm sorry to hear that. Would you like to talk about it?"} and extracts the text portion.

[1129] Step 12:

[1130] Convert text received by the device into speech. Use a speech synthesis API to convert text into speech data. Example: Use the following API:

[1131] python

[1132] synthesized_audio = text_to_speech("I'm sorry to hear that. Would you like to talk about it?")

[1133] Step 13:

[1134] The terminal plays the generated voice data to the user, and provides the voice to the user through the terminal's speaker.

[1135] This allows the user to receive an appropriate voice response that corresponds to their own emotions.

[1136] Example 2

[1137] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1138] Conventional communication support systems have a limited ability to properly analyze text entered by users and provide responses based on that content. Furthermore, they are unable to generate responses that reflect the user's emotions, making it difficult to improve the quality of communication. Furthermore, functions such as text translation and summarization are often provided separately, making them less convenient as an integrated system.

[1139] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for a user to input text, means for receiving the text, shaping it into a data format, and then transmitting it to the server for analysis, means for the server to analyze the text and determine a task, means for the generative AI model and emotion engine to perform appropriate processing based on the task, means for converting the generated response into speech, and means for providing the speech to the user. This makes it possible to automatically provide appropriate translations, summaries, suggestions for creative ideas, and even responses according to emotions based on the text entered by the user.

[1140] "User" refers to any individual or corporation that uses this system.

[1141] "Means for inputting text" refers to an interface for a user to input text information, such as a keyboard or a touch screen.

[1142] "Data format" refers to a format for formatting information based on certain rules, and includes, for example, JSON format.

[1143] A "server" refers to a computer system that provides services to other computers over a network.

[1144] A "generative AI model" refers to an artificial intelligence model that is trained using large amounts of text data and generates and analyzes text.

[1145] "Emotion engine" refers to a system for analyzing a user's emotions from input text.

[1146] "Means for converting a response into voice" refers to technology for converting text data into voice data, including, for example, a voice synthesis API.

[1147] A "task" refers to a specific process or function that the system must perform, such as translation, summarization, creative idea suggestion, or sentiment analysis.

[1148] "Analysis" refers to the process of analyzing input text data and clarifying its content and intent.

[1149] This invention relates to a communication support system that performs advanced processing of text entered by a user, and provides translation, summarization, creative idea suggestions, and even responses tailored to the user's emotions. This system consists of three main components: the user, the terminal, and the server. Each component and its operation are described in detail below.

[1150] User operations and device roles

[1151] A user inputs text through a terminal such as a personal computer or smartphone. The terminal is provided with a text input field, and when the user inputs text, the text is received by the terminal. The input text is formatted into a standard data format such as JSON. This formatted data is then sent to the server via an HTTP request.

[1152] Data analysis and model execution on the server

[1153] The server analyzes the received data and determines whether it corresponds to a translation task, a summarization task, a creative idea suggestion, or sentiment analysis. This determination process uses a text analysis library (e.g., NLTK or spaCy). After determining the specific task, the server invokes a generative AI model and sentiment engine to perform the necessary processing.

[1154] Generative AI models are pre-trained using large amounts of text data, enabling accurate analysis and response generation. For example, OpenAI's GPT-3 is used. Meanwhile, emotion engines are systems that recognize user emotions from input text, such as IBM Watson's Tone Analyzer.

[1155] Response generation and information return

[1156] The response text generated by the generative AI model is again formatted in JSON format. This response data is sent from the server to the device. The device then uses speech synthesis technology to convert the received response text into audio data. For example, Google Cloud Text-to-Speech API is used. The converted audio data is then provided to the user via the speaker.

[1157] Specific examples

[1158] Below is a concrete example of a prompt entered by a user and the system's response to it.

[1159] Example 1: Translation task

[1160] When a user types "Translate 'Hello' to Japanese," the device sends the text to the server, and the generative AI model generates the translation result "Hello," which is then converted into speech on the device and played back to the user.

[1161] Example 2: Emotional response task

[1162] If a user types "I had a bad day today," the emotion engine recognizes the emotion "negative," and the generative AI model generates the corresponding message, "I'm sorry to hear that. Would you like to talk about it?" This message is also converted to audio and provided to the user.

[1163] Example 3: Creative idea proposal task

[1164] If a user types "Give me an idea for a science project," the generative AI model generates the idea "Build a small wind turbine using household materials," which is also provided to the user via voice synthesis.

[1165] In this way, each component works in cooperation with the others to efficiently and flexibly provide a variety of responses according to the text entered by the user.

[1166] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1167] Program processing flow and specific explanation

[1168] Step 1:

[1169] The user inputs text. For example, the user inputs "Translate 'Hello' to Japanese" into a text input field on the terminal. The input text is received by the terminal.

[1170] Step 2:

[1171] The device formats the received text into JSON format. Specifically, it uses JavaScript to convert the input text into an object format, and then uses the JSON.stringify function to convert it into a JSON string. This formatted data is sent as an HTTP POST request to the API endpoint.

[1172] Input: Raw text entered by the user ("Translate 'Hello' to Japanese")

[1173] Output: JSON formatted data

[1174] Step 3:

[1175] The server receives the request and parses the received JSON data. The server receives the request using the Python Flask framework, and parses the JSON data using request.get_json().

[1176] Input: JSON format data

[1177] Output: Parsed text ("Translate 'Hello' to Japanese")

[1178] Step 4:

[1179] The server analyzes the parsed text to determine whether it is a translation, summary, creative idea suggestion, or sentiment analysis, for example using a text analysis library (NLTK or spaCy).

[1180] Input: Parsed text

[1181] Output: The type of task (in this case, a translation task)

[1182] Step 5:

[1183] The server calls a generative AI model or emotion engine based on the identified task. For example, for a translation task, it sends a translation request to a generative AI model (e.g., GPT-3). In this case, it uses the GPT-3 API to send a request to translate "Hello" into Japanese.

[1184] Input: Task type (translation task) and source text ("Hello")

[1185] Output: Translation result("Hello")

[1186] Step 6:

[1187] The generated response text is then formatted into JSON again. The server then uses the jsonify function to convert the translation result into JSON format and sends it back as an HTTP response.

[1188] Input: Translation result("Hello")

[1189] Output: Response data in JSON format

[1190] Step 7:

[1191] The device analyzes the received response data and converts it into audio. The device uses JavaScript to parse the JSON data and converts the translation results into audio data using a speech synthesis API (such as Google Cloud Text-to-Speech). The generated audio data is played using the browser's Audio object.

[1192] Input: Response data in JSON format

[1193] Output: Audio data

[1194] Step 8:

[1195] The device plays the generated audio data through the speaker and provides it to the user as sound. Specifically, it plays the audio using the Audio.play() method.

[1196] Input: Audio data

[1197] Output: Speech provided to the user ("Hello")

[1198] The above is the specific processing flow of this system and the detailed operation of each step, which allows users to efficiently and flexibly obtain results for translation and other tasks.

[1199] (Application example 2)

[1200] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1201] Conventional food delivery services require users to input their menu items through unintuitive text input, and are unable to provide appropriate suggestions based on the user's emotions. This results in a poor user experience and reduces the appeal of the service.

[1202] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input text, means for receiving the text and transferring it to a generative AI model for analysis, means for the generative AI model to generate an appropriate response based on the text, means for converting the generated response into speech, means for providing the speech to the user, means for analyzing the user's emotions using the generative AI model and an emotion engine, and means for suggesting an appropriate menu to the user based on the emotion analysis. This allows the user to intuitively place an order using natural language and receive appropriate menu suggestions based on their emotions.

[1203] A "means for user text input" is a device or software that provides an interface for a user to input text in a natural language.

[1204] The "means for receiving said text and forwarding it to a generative AI model for analysis" refers to a function that sends user-entered text to a server and prepares the data to be analyzed by a generative AI model.

[1205] The "means by which the generative AI model generates an appropriate response based on the text" refers to the algorithm or process by which the generative AI model generates an appropriate reply or suggestion based on the text analyzed by the generative AI model.

[1206] "Means for converting the generated response into speech" refers to a technology or system that converts the text response generated by the generative AI model into speech data.

[1207] The "means for providing the audio to the user" refers to a device such as a speaker or earphone for playing back the data converted into audio to the user.

[1208] The "means for analyzing user emotions using the generative AI model and emotion engine" refers to a system that combines a generative AI model and an emotion analysis engine to identify emotions from the user's input text and adjust responses and suggestions based on the results.

[1209] The "means for suggesting an appropriate menu to a user based on the emotion analysis" is a system that automatically suggests an appropriate menu to a user in a food delivery service based on the user's emotions analyzed by an emotion analysis engine.

[1210] This invention is a system that processes user-entered text using a generative AI model and an emotion engine to support ordering and suggestions in food delivery services. This system consists of a user, a terminal, and a server, and achieves high efficiency and flexibility through the collaboration of each component.

[1211] A user inputs text through their device. The device has a text input field, and when the user types, the text is received by the device. This text is then formatted into a standard data format such as JSON and sent to the server.

[1212] The server analyzes the data received from the device and determines whether it includes a translation task, a summary task, creative suggestions, or sentiment analysis. The generative AI model and the emotion engine work together to execute a specific task. Specifically, the server sends an analysis request to the generative AI model to request the necessary processing. The emotion engine also analyzes the user's sentiment from the input text, and the generative AI model adjusts its response accordingly.

[1213] The generative AI model is pre-trained using a large amount of text data, enabling highly accurate analysis and response generation. The emotion engine can recognize the user's emotions from a variety of input data, including text, voice, and facial expressions. For example, if the text entered by the user contains emotions such as "I had a tough day, recommend me something soothing," the emotion engine will identify emotions such as fatigue and anxiety, and the generative AI model will generate an appropriate response based on this.

[1214] The response text generated by the generative AI model is formatted in JSON and returned to the server. The server receives this generated response text and sends it back to the device. The device then uses speech synthesis technology to convert the received text into audio data. The text is converted into audio through a speech synthesis API, and the audio data is provided to the user through the speaker.

[1215] For example, if a user types "I want to order something spicy," the generative AI model will suggest "spicy dish suggestions," which will be received by the device and played aloud to the user via speech synthesis.

[1216] For example, if a user types, "I had a tough day, recommend me something soothing," the emotion engine will recognize emotions such as "fatigue" and "anxiety," and the generative AI model will generate an appropriate response, such as "recommended smoothies to relax you."

[1217] In this way, the system of the present invention allows users to efficiently obtain translations, summaries, and creative suggestions, as well as receive appropriate menu suggestions based on their current emotions, thereby improving the user experience in food delivery services and increasing their appeal.

[1218] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1219] Step 1:

[1220] The user inputs text into the terminal. The input text includes the user's order details or questions. For example, "I want to order something spicy." The input is formatted into a standard data format such as JSON. The input data is sent from the terminal to the server.

[1221] Step 2:

[1222] The server receives the text sent from the device, stores it in a database, and formats it appropriately for analysis. The server analyzes the text and determines whether it needs to be translated, summarized, suggested, or sentiment analyzed.

[1223] Step 3:

[1224] The server sends an analysis request to the generative AI model, which generates a highly accurate response based on the received text. For example, if a user types "I want to order something spicy," the generative AI model generates "spicy dish suggestions."

[1225] Step 4:

[1226] The server receives the response from the generative AI model and passes the text to the emotion engine to analyze the user's emotions. The emotion engine identifies the emotions contained in the text and adjusts the response accordingly. For example, for the input "I had a tough day, recommend me something soothing," the emotion engine identifies emotions such as "fatigue" and "anxiety."

[1227] Step 5:

[1228] The server combines the generated response text with the sentiment analysis results to generate optimal suggestions and responses, and the combined results are formatted in JSON format.

[1229] Step 6:

[1230] The server sends the formatted response text to the device, which then uses speech synthesis technology to provide a response to the user. The speech synthesis API converts the text into speech and provides it to the user as audio data through the speaker.

[1231] Step 7:

[1232] Users receive responses and suggestions through voice data, such as suggestions for spicy dishes or relaxing menu items, based on their emotions and the order they make.

[1233] In this way, a system is constructed that uses a generative AI model and emotion engine to perform advanced analysis and data calculations based on the user's input at each processing step, providing the user with an appropriate response.

[1234] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1235] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1236] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1237] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1238] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1239] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1240] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1241] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1242] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1243] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1244] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1245] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1246] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1247] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1248] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1249] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1250] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1251] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1252] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1253] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1254] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1255] The following is further disclosed regarding the above embodiment.

[1256] (Claim 1)

[1257] a means for a user to input text;

[1258] means for receiving the text and forwarding it to a generative AI model for analysis;

[1259] means for the generative AI model to generate an appropriate response based on the text;

[1260] means for converting the generated response into speech;

[1261] means for providing said audio to a user;

[1262] A system including:

[1263] (Claim 2)

[1264] 10. The system of claim 1, wherein the generative AI model includes means for performing a translation of the text.

[1265] (Claim 3)

[1266] 10. The system of claim 1, wherein the generative AI model includes means for summarizing the text.

[1267] "Example 1"

[1268] (Claim 1)

[1269] a means for a user to input text;

[1270] means for receiving and formatting said text into a standard data format;

[1271] means for transmitting the formatted data to a server;

[1272] A means for analyzing the data received by the server and requesting processing from a generative AI model;

[1273] means for the generative AI model to generate an appropriate response based on the text;

[1274] means for receiving the generated response and transmitting it to a terminal;

[1275] means for converting the received response into audio data;

[1276] means for providing said audio data to a user;

[1277] A system including:

[1278] (Claim 2)

[1279] 10. The system of claim 1, wherein the generative AI model includes means for performing a translation of the text.

[1280] (Claim 3)

[1281] 10. The system of claim 1, wherein the generative AI model includes means for summarizing the text.

[1282] "Application Example 1"

[1283] (Claim 1)

[1284] a means for a user to input text;

[1285] means for receiving the text and forwarding it to a generative AI model for analysis;

[1286] means for the generative AI model to generate an appropriate response based on the text;

[1287] means for converting the generated response into speech;

[1288] means for providing said audio to a user;

[1289] The system is installed on a factory robot, and a means for generating a summary of the work procedure and work status using a generative AI model in response to instructions from a worker and providing the summary in audio form;

[1290] A system including:

[1291] (Claim 2)

[1292] 10. The system of claim 1, wherein the generative AI model includes means for performing a translation of the text.

[1293] (Claim 3)

[1294] 10. The system of claim 1, wherein the generative AI model includes means for summarizing the text.

[1295] "Example 2: Combining Emotion Engines"

[1296] (Claim 1)

[1297] a means for a user to input text;

[1298] means for receiving the text, formatting it into a data format, and then transmitting it to a server for analysis;

[1299] means for the server to analyze the text and determine a task;

[1300] means for the generative AI model and emotion engine to take appropriate action based on the task;

[1301] means for converting the generated response into speech;

[1302] means for providing said audio to a user;

[1303] A system including:

[1304] (Claim 2)

[1305] 10. The system of claim 1, wherein the generative AI model performs translation of the text.

[1306] (Claim 3)

[1307] 10. The system of claim 1, wherein the generative AI model summarizes the text.

[1308] "Application example 2 when combining emotion engines"

[1309] (Claim 1)

[1310] a means for a user to input text;

[1311] means for receiving the text and forwarding it to a generative AI model for analysis;

[1312] means for the generative AI model to generate an appropriate response based on the text;

[1313] means for converting the generated response into speech;

[1314] means for providing said audio to a user;

[1315] A means for analyzing a user's emotions using the generative AI model and an emotion engine;

[1316] means for suggesting an appropriate menu to a user based on the emotion analysis;

[1317] A system including:

[1318] (Claim 2)

[1319] 10. The system of claim 1, wherein the generative AI model includes means for performing a translation of the text.

[1320] (Claim 3)

[1321] 10. The system of claim 1, wherein the generative AI model includes means for summarizing the text. [Explanation of symbols]

[1322] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for a user to input text; means for receiving the text and forwarding it to a generative AI model for analysis; means for the generative AI model to generate an appropriate response based on the text; means for converting the generated response into speech; means for providing said audio to a user; A system including:

2. The system of claim 1 , wherein the generative AI model includes means for performing a translation of the text.

3. The system of claim 1 , wherein the generative AI model includes means for summarizing the text.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A