System

A system with a server and terminal device provides real-time translation and emotion recognition, addressing language barriers by ensuring accurate and emotionally sensitive communication in international contexts.

JP2026019086APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120495
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Language barriers pose a significant obstacle to smooth international communication and business interactions, particularly in regions with a limited number of fluent foreign language speakers, necessitating accurate and instantaneous translation solutions.

Method used

A system that includes a server for receiving and analyzing user input, translating it in real-time using a generative AI model, and providing the translation results to a terminal device, equipped with error handling and emotion recognition capabilities to enhance communication accuracy and emotional reflection.

Benefits of technology

Enables high-accuracy, real-time translation across languages, facilitating smooth communication and improving service quality by reflecting user emotions, particularly beneficial in business and tourism sectors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019086000001_ABST
    Figure 2026019086000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: means for receiving input from a user; means for parsing the received input; means for translating the parsed input into a specified language; and means for providing the translated result to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] With the advancement of international exchange and business in modern society, multilingual communication is becoming increasingly important. However, language barriers are a major obstacle to such exchange. In Japan in particular, the small number of people who speak foreign languages ​​fluently makes it difficult to communicate smoothly with business people and tourists from abroad. In this context, there is a growing demand for systems that can provide accurate and instantaneous translation in real time. The purpose of this invention is to eliminate communication difficulties caused by language barriers and facilitate international exchange and business. [Means for solving the problem]

[0005] The present invention solves the above problems by the following means. Specifically, it provides a system including a means for receiving input from a user, a means for analyzing the received input, a means for translating the analyzed input into a specified language, and a means for providing the user with the translated result. The system considers whether the user's input is text or speech and uses a generative model to translate the input content with high accuracy and in real time. It also has a function for analyzing the user's ID and target language information, providing translations optimized for each individual user. Furthermore, the system is used in the business and tourism fields, and is equipped with an error handling means for performing error checks at each processing step, providing a highly reliable translation service. This breaks down language barriers and enables smooth communication in the international business and tourism fields.

[0006] A "user" is a person or entity that provides input to utilize the system.

[0007] "Input" refers to text or voice data provided by a user to a system.

[0008] A "terminal" is a device through which a user makes input and is a device that communicates with the system.

[0009] "Server" refers to the central system that receives input from users and performs the parsing and interpretation process.

[0010] "Receiving means" refers to a function that allows the server to receive user input.

[0011] "Analysis means" refers to a function for interpreting and preprocessing received user input data.

[0012] "Translation means" refers to a function for converting analyzed input data into a specified language.

[0013] "Providing means" refers to a function for displaying or outputting the translated results to the user.

[0014] "Generative model" refers to the AI ​​technology used to translate user input with high accuracy and in real time.

[0015] "Error handling means" refers to a function that checks for errors at each step of processing and notifies an error message if necessary. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention is a system that translates text and speech input by a user into other languages ​​in real time, facilitating communication. This system is mainly composed of three entities: a server, a terminal, and a user.

[0038] overview

[0039] A user accesses the system and enters text or voice into an input field. For example, a user might enter "Hello, what time is it?"

[0040] The device receives the user's input and sends it to the server, which analyzes the received data and translates it into a specified language (e.g., English) using a generative model.

[0041] Once the translation is complete, the server sends the translation results to the terminal, which then displays them to the user, allowing users to communicate smoothly with users of other languages.

[0042] Specific explanation of each element

[0043] User Input

[0044] Users access the platform through their own devices and enter text or voice into the input fields, which can be used in a variety of situations, such as business meetings or exchanging information at tourist spots.

[0045] For example, consider the case where a user uses their smartphone to type "Hello, what time is it now?" in Japanese.

[0046] Device operation

[0047] The device receives the user's input in real time and transmits the data to the server, for example via its own HTTP protocol, including the user's input text, user ID, and desired language information.

[0048] Server Processing

[0049] The server analyzes the received data. Once the analysis is complete, it uses a generative model to perform the translation process. The generative model utilizes the latest AI technology to achieve high accuracy and real-time translation.

[0050] For example, the server parses the Japanese phrase "Hello, what time is it now?" and translates it into English using the appropriate model, resulting in the translation "Hello, what time is it now?"

[0051] Providing translation results

[0052] After the translation is complete, the server sends the results to the terminal, which then displays them to the user in a user-friendly format, making them instantly available.

[0053] Specific examples

[0054] As a concrete example, the translation flow from Japanese to English is shown below.

[0055] The user types, "Hello, what time is it?"

[0056] The terminal sends the input to the server.

[0057] The server parses the received input and translates it into "Hello, what time is it now?"

[0058] The server sends the translation results to the terminal.

[0059] The terminal displays the translation results to the user.

[0060] This allows users to communicate smoothly with speakers of other languages ​​in real time. This system is particularly expected to be used in the business and tourism sectors, and its error handling functionality will provide a highly reliable service.

[0061] The processing flow will be explained below.

[0062] Step 1:

[0063] A user accesses the platform and enters text or voice into an input field, for example, "Hello, what time is it?"

[0064] Step 2:

[0065] The device receives the user's input and sends the data to the server, including the user's input text, user ID, and desired language information.

[0066] Step 3:

[0067] The server then analyzes the received data, which includes identifying the language of the text entered by the user and performing any necessary pre-processing (e.g., noise removal and text normalization).

[0068] Step 4:

[0069] The server uses the generative model to translate the parsed input data into the specified language. For example, to translate the Japanese phrase "Hello, what time is it now?" into English, it would be converted to "Hello, what time is it now?"

[0070] Step 5:

[0071] The server formats the translation results and sends them to the device, ensuring that the formatted data is organized in a way that is easy for the user to understand.

[0072] Step 6:

[0073] The device receives the translation and displays it to the user, properly formatted in a user-friendly language.

[0074] Step 7:

[0075] The user checks the displayed translation result and takes the next action (re-entering or other operation). If there is further input, the cycle starts again from step 1.

[0076] Through this series of processes, users can communicate smoothly and in real time in other languages.

[0077] Example 1

[0078] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0079] Real-time communication between multiple languages ​​is important, especially in business and tourism. However, conventional translation systems often suffer from low translation accuracy and lack real-time capabilities, making them prone to errors. To address this issue, a system that provides high-accuracy, real-time translation and is easy for users to use is needed.

[0080] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0081] In this invention, the server includes means for parsing received input, means for using a generative AI model to translate the parsed input into a specified language, and means for providing a prompt sentence to the generative AI model, thereby enabling highly accurate and real-time translation.

[0082] "Input" is text or voice data that a user provides to the system via a terminal.

[0083] "Analysis" is the process by which the server understands the input data it receives and extracts the necessary information.

[0084] A "generative AI model" is an artificial intelligence model that uses technologies such as neural networks to translate input data into a specified language.

[0085] A "prompt sentence" is a sentence that instructs the generative AI model on what processing to do.

[0086] "Translation" is the process of converting input data into a specified language.

[0087] "Providing" is the process in which the server sends the translation result to the terminal, and the terminal displays the result to the user.

[0088] A "terminal" is a device (e.g., a smartphone or tablet) that a user uses to access the system and input data or display results.

[0089] "Server" means a central processing unit that receives input from a terminal and performs analysis, translation, and provides results.

[0090] "Result display" is the process by which the terminal visually presents the translation results received from the server to the user.

[0091] The present invention provides a system for translating text and speech input by a user into other languages ​​in real time to facilitate communication. The system operates in cooperation with a user, a terminal, and a server. A specific embodiment of the system is described below.

[0092] User Input

[0093] Users access the system using devices such as smartphones or tablets. They enter text into the provided input fields or use a microphone to input voice. For example, consider a case where a user enters "Hello, what time is it now?" in Japanese.

[0094] Device operation

[0095] The device receives user input in real time and sends the data to the server. The device formats the user input data and adds necessary metadata such as the user ID and the language to be translated. This data is sent to the server via the HTTP protocol. For example, the device generates an HTTP request containing the text data "Hello, what time is it now?" and sends it to the server.

[0096] Server Processing

[0097] The server receives the request sent from the device and extracts the input text and necessary metadata from the request body. After verifying that the received data is free of errors, it performs the translation using a generative AI model. Specifically, the server generates a prompt sentence for the generative AI model (e.g., OpenAI's GPT-4) and passes it on. An example of a prompt sentence is, "Please translate from Japanese to English. Input sentence: 'Hello, what time is it now?'"

[0098] Providing translation results

[0099] Once the translation process is complete, the server sends the translation result to the terminal as an HTTP response. The response includes the translated text (e.g., "Hello, what time is it now?"). At this time, the response is also checked for formatting and errors.

[0100] Displaying results on your device

[0101] The device analyzes the response received from the server and displays the results to the user. For example, the translation result "Hello, what time is it now?" is displayed on the device's display. The user can check the result in real time.

[0102] Specific examples

[0103] Below is a concrete Japanese to English translation flow:

[0104] 1. The user types "Hello, what time is it?" into the text field.

[0105] 2. The device sends the input to the server along with the "User ID" and "Translation Language: English."

[0106] 3. The server receives the data "Hello, what time is it now?" and checks the content.

[0107] 4. The server sends the generative AI model a prompt: "Please translate from Japanese to English. Input sentence: 'Hello, what time is it now?'" and receives a response from the generative AI model: "Hello, what time is it now?"

[0108] 5. The server sends the translation result, "Hello, what time is it now?" to the terminal.

[0109] 6. The terminal displays the result to the user: "Hello, what time is it now?"

[0110] This system will enable users to communicate smoothly with speakers of other languages ​​in real time, and is expected to be particularly useful in the fields of business and tourism.

[0111] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0112] Step 1: User Input

[0113] A user accesses the system through his / her own terminal. The user inputs text or voice into an input field. Specifically, if the user wants to input "Hello, what time is it now?" in Japanese, he / she types the characters into the input field of the terminal or captures the voice using the microphone. The input data (text or voice) is prepared.

[0114] Step 2: Send data from the device to the server

[0115] The terminal formats the user's input data and adds the necessary metadata (user ID, language to translate, etc.), then sends the formatted request to the server using the HTTP protocol. The input data and metadata are sent to the server.

[0116] Step 3: Data analysis on the server

[0117] The server receives the request from the terminal and analyzes its contents. Specifically, it extracts the input text, user ID, and desired language information from the request body. The server then verifies that the data format and content are correct and performs error checking. The analysis results in structured data.

[0118] Step 4: Translation on the server

[0119] The server performs translation using a generative AI model based on the analysis results. Specifically, it generates a prompt sentence for the generative AI model (e.g., generative model GPT-4) and passes it on. For example, a prompt sentence such as "Please translate from Japanese to English. Input sentence: 'Hello, what time is it now?'" is generated and input into the AI ​​model. The generative AI model converts the input data into the specified language and outputs the translated data, "Hello, what time is it now?"

[0120] Step 5: Send the translation results from the server to the device

[0121] The server formats the translation results obtained from the generative AI model as an HTTP response and sends it to the device. The response includes the translation result (e.g., "Hello, what time is it now?"). The server checks the response format and content for errors. The formatted response data is sent to the device.

[0122] Step 6: View the results in your device

[0123] The device analyzes the response received from the server and extracts the translation result. Specifically, the device extracts the translated text from the response data and displays it to the user. The displayed result (e.g., "Hello, what time is it now?") is confirmed by the user.

[0124] Through these steps, users can obtain highly accurate translation results into other languages ​​in real time.

[0125] (Application example 1)

[0126] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0127] In modern society, situations requiring multilingual communication in brick-and-mortar stores are on the rise, but language barriers can make smooth customer service difficult. This problem is particularly pronounced in regions where the number of foreign language speakers is increasing, requiring accurate translation in real time. However, existing translation systems often lack sufficient translation speed and accuracy. Furthermore, they lack the ability to check history, making it impossible to refer to past communications, which risks reducing the quality of service.

[0128] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0129] In this invention, the server includes means for receiving input from a user, means for analyzing the received input, means for translating the analyzed input into a specified language, means for providing the translated result to the user, and means for providing an interface for users and foreign language speakers to communicate in real time at the store. This enables accurate multilingual communication in real time. The inclusion of a history function also allows past communications to be referenced, improving the quality of service. Furthermore, translation is performed using a generative model, and accuracy is improved by specific prompt sentences, providing more reliable translation results.

[0130] "User" refers to a person who uses the system to input text or voice and receive the results translated into another language.

[0131] "Input" refers to the text or voice information provided by a user to a system.

[0132] "Analysis" refers to the information processing means that processes received input data and makes sense of it.

[0133] "Translation" refers to the process of converting parsed input data into another specified language.

[0134] "Interface" refers to the operation screen and operation means that facilitate smooth interaction between the user and the system.

[0135] The "history function" refers to a function that saves past communication details with users and allows them to refer to them as needed.

[0136] "Generative model" refers to a machine learning algorithm that uses AI technology to produce highly accurate translations.

[0137] A "prompt sentence" refers to an instruction sentence given to a generative model to perform an appropriate translation.

[0138] To implement this invention, it is necessary to build a system in which users, terminals, and servers work together to provide highly accurate real-time translation and support smooth communication between multiple languages.

[0139] User operations

[0140] Users access the system using devices such as smartphones or tablets. The device displays an interface as an operation screen, where users input text or voice. For example, consider the case where a store clerk asks a foreign customer, "Hello, how do you want to use this product?" in Japanese.

[0141] Device operation

[0142] The terminal receives text and voice input from the user and sends it to the server, generating a data packet containing the user's input data, user ID, and desired language information using a communication method such as the HTTP protocol.

[0143] Server Processing

[0144] The server analyzes the input data sent from the device. After analysis, it uses a generative AI model to translate it into the specified language. Specifically, it uses OpenAI's API to translate based on the following prompt:

[0145] Prompt Sentence Examples

[0146] "Translate this text to English: Hello, how do I use this product?"

[0147] Based on this prompt, the generative model generates a translation result such as "Hello, how do I use this product?"

[0148] Providing translation results

[0149] The server then sends the translation result back to the terminal, which then displays it in a format that is easy for the user to understand. For example, the clerk's terminal screen might display "Hello, how do I use this product?" in English.

[0150] History function

[0151] Furthermore, the system has a built-in history function that allows past communications to be saved, allowing users to refer to past conversation history as needed to improve the quality of service.

[0152] Hardware and Software

[0153] Hardware: Smartphone or tablet (iOS or Android)

[0154] Software: The front end uses HTML, CSS, and JavaScript, and the back end uses Python Flask and the OpenAI API.

[0155] By combining these steps, the system facilitates multilingual communication in brick-and-mortar stores, enabling users and foreign language speakers to communicate effectively in real time.

[0156] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0157] Step 1:

[0158] A user accesses the system using a device such as a smartphone or tablet and inputs text or voice into an input field. For example, a store clerk might input "Hello, how do you want to use this product?" in Japanese. In this case, the input is sent to the device as text or voice data.

[0159] Step 2:

[0160] The device analyzes the received text and voice data and sends it to the server. At this point, a data packet is formed using the HTTP protocol or similar, and information such as the user ID and desired language (e.g., English) is also included. This is the process by which the text and voice data received as input is converted into a data packet to be sent to the server.

[0161] Step 3:

[0162] The server receives data packets sent from the device and analyzes the input data. After analysis, it uses a generative AI model to translate it into the specified language. For example, the server provides the prompt sentence "Translate this text to English: Hello, how do you use this product?" to the generative model, and the resulting translation is "Hello, how do I use this product?"

[0163] Step 4:

[0164] The server then sends the translation results back to the terminal. This time, the translation results are sent in text format. The server's input is the translated text data, and it is the process that processes it to send it to the terminal.

[0165] Step 5:

[0166] The terminal receives the translation results from the server and displays them to the user. The results are displayed visually through a user-friendly interface, allowing the user to understand them immediately. For example, the salesperson's terminal screen might display "Hello, how do I use this product?"

[0167] Step 6:

[0168] As a history function, the device will store past translation history and allow users to refer to this history as needed, supporting continuous communication with the same user and improving the quality of service. The history data will be stored in local storage or cloud storage, allowing users to access it at any time.

[0169] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0170] This invention is a system that translates user-entered text and speech into other languages ​​in real time, and also recognizes the user's emotions and reflects them in the translation results, facilitating smooth communication. This system is mainly composed of four main components: a server, a terminal, a user, and an emotion engine.

[0171] overview

[0172] A user accesses the system and inputs text or voice into an input field. For example, consider the case where a user inputs "Hello, what time is it now?". At the same time, the system reads the user's emotions from their tone of voice and facial expressions.

[0173] The device receives user input and sends the data and emotion information to the server, which analyzes the received data and translates it into a specified language (e.g., English) using a generative model. It also uses an emotion engine to recognize the user's emotion and adjust the translation results based on that emotion.

[0174] Once the translation is complete, the server sends the translation results to the terminal, which then displays them to the user, allowing users to communicate smoothly with users of other languages.

[0175] Specific explanation of each element

[0176] User Input

[0177] Users access the platform through their own devices and enter text or voice into the input field. This is used in a variety of situations, such as business meetings and exchanging information at tourist spots. For example, consider a case where a user uses their smartphone to enter "Hello, what time is it now?" in Japanese.

[0178] Device operation

[0179] The device receives user input in real time and transmits the data and emotion information to the server via its own HTTP protocol, etc. The transmitted data includes the user's input text, user ID, desired language information, and emotion information.

[0180] Server Processing

[0181] The server analyzes the received data. Once the analysis is complete, it uses a generative model to perform the translation process. The generative model utilizes the latest AI technology to achieve high accuracy and real-time translation.

[0182] For example, the server analyzes the Japanese phrase "Hello, what time is it now?" and translates it into English using an appropriate model, resulting in the translation "Hello, what time is it now?"

[0183] In parallel, the server uses an emotion engine to recognize the user's emotions and adjusts the translated text accordingly. For example, if the user is excited, the server might translate it as "What time is it right now?"

[0184] Providing translation results

[0185] After the translation is complete, the server sends the results to the device, which then displays them to the user in a user-friendly format, taking into account emotional information.

[0186] Specific examples

[0187] As a concrete example, the translation flow from Japanese to English is shown below.

[0188] A user types, "Hello, what time is it?" At the same time, the user's tone of voice is analyzed and it is detected that the user is excited.

[0189] The device sends the input and emotion information to the server.

[0190] The server analyzes the received input and translates it as "Hello, what time is it now?". It also adjusts the translation result based on the emotion information to "What time is it right now?"

[0191] The server sends the translation results to the terminal.

[0192] The terminal displays the translation results to the user.

[0193] This allows users to communicate smoothly and emotionally with speakers of other languages ​​in real time. This system is particularly expected to be used in the business and tourism sectors, and its error handling functionality will provide a highly reliable service.

[0194] The processing flow will be explained below.

[0195] Step 1:

[0196] A user accesses the platform and enters text or voice into an input field, for example, "Hello, what time is it?" The user's facial expressions and tone of voice are also recorded.

[0197] Step 2:

[0198] The device receives user input and converts it into text data, while also collecting the user's emotional information (e.g., excitement, sadness, joy).

[0199] Step 3:

[0200] The device sends the user's input data and emotion information to the server, along with the user ID and desired language.

[0201] Step 4:

[0202] The server analyzes the received data, which includes identifying the language of the text and pre-processing it (denoising, normalization).

[0203] Step 5:

[0204] The server uses the generative model to translate the parsed text into the specified language. For example, the Japanese phrase "Hello, what time is it now?" is translated into English as "Hello, what time is it now?"

[0205] Step 6:

[0206] The server uses the emotion engine to recognize the user's emotion. For example, if the user is excited, the emotion is recognized as "excited."

[0207] Step 7:

[0208] The server adjusts the translation results based on the emotional information it recognizes. For example, if the user is feeling excited, it will modify "Hello, what time is it now?" to reflect that emotion, such as "What time is it right now?"

[0209] Step 8:

[0210] The server formats the final translation and sends it to the device, with adjustments based on emotional information.

[0211] Step 9:

[0212] The device will then display the translation results to the user, specifically asking, "What time is it right now?"

[0213] Step 10:

[0214] The user checks the displayed translation result and takes the next action (re-entering or other operations) if necessary. If re-entering is required, the cycle starts again from step 1.

[0215] Through this series of processes, users can communicate smoothly and emotionally with speakers of other languages ​​in real time.

[0216] Example 2

[0217] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0218] In modern society, real-time communication between people who speak different languages ​​is becoming increasingly important. However, conventional translation systems simply translate text and are unable to provide translation results that reflect the user's emotions. This makes it difficult to accurately convey important emotional information in communication.

[0219] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving input from a user, means for analyzing the received input, means for translating the analyzed input into a specified language, means for recognizing the user's emotion, means for adjusting the translation result based on the emotion recognition result, and means for providing the translation result to the user. This makes it possible to provide a translation result that reflects emotional information contained in the text or voice input by the user.

[0220] "User" means any individual or entity that uses the System.

[0221] "Terminal" refers to an electronic device that allows a user to access and operate the system. Examples include smartphones and personal computers.

[0222] A "server" is a computer system that receives data sent from a terminal and analyzes and processes the data.

[0223] "Means for receiving input" refers to a mechanism by which the terminal can receive text or voice data entered by a user.

[0224] "Means for analyzing input" refers to a mechanism for analyzing received data and understanding its meaning and intent. This includes natural language processing technology.

[0225] "Means for translating" means a mechanism for translating analyzed text into another specified language, including a generative AI model.

[0226] "Means for recognizing emotions" refers to a mechanism for detecting emotions from the user's tone of voice and facial expressions and acquiring that information.

[0227] The "means for adjusting the translation result" refers to a mechanism for correcting the translated text based on the recognized emotional information and adding appropriate emotional expressions.

[0228] "Means for providing the translated results to the user" refers to a mechanism for displaying or audibly notifying the user of the final translation results via a terminal.

[0229] A "generative AI model" is an algorithm that uses artificial intelligence techniques to generate or process data. For example, it includes deep learning models that specialize in natural language processing.

[0230] A "prompt sentence" refers to the text data input into a generative AI model, and is the initial information that the model uses to translate and generate.

[0231] This system translates user-input text and speech into other languages ​​in real time, and also recognizes the user's emotions and reflects them in the translation results. This system is mainly composed of four main components: a server, a terminal, a user, and an emotion recognition engine.

[0232] System configuration and technologies used

[0233] User

[0234] A user accesses their device and inputs text or voice into an input field. For example, a user might input "Hello, what time is it now?" in Japanese. The device then sends the user's tone of voice and facial expressions to an emotion recognition engine.

[0235] Terminal

[0236] The device receives user input in real time and transmits the data and emotion information to the server. The device can be a smartphone or a PC, and data transmission uses the HTTP protocol. The transmitted data includes the input text, user ID, desired language information, and emotion information.

[0237] server

[0238] The server analyzes the data received from the device and translates it into the specified language using a generative AI model. Generative AI models use algorithms specialized for natural language processing, such as OpenAI's GPT-3 model. The server uses an emotion recognition engine (such as Microsoft Azure Cognitive Services) to recognize the user's emotions and adjust the translation results based on that information.

[0239] Emotion Recognition Engine

[0240] The emotion recognition engine analyzes the user's tone of voice and facial expression data to obtain emotional information. Based on this emotional information, the server adjusts the translation results and adds appropriate emotional expressions.

[0241] Specific operation flow

[0242] 1. The user types "Hello, what time is it?". At the same time, the user's excitement level is collected as emotional information.

[0243] 2. The device sends the input and emotion information to the server.

[0244] 3. The server receives the request and parses the data.

[0245] 4. The server uses the generative AI model to translate it as "Hello, what time is it now?"

[0246] 5. The server sends the data to the emotion recognition engine to check the user's excitement level.

[0247] 6. The server adjusts the translation result to "What time is it right now?"

[0248] 7. The server sends the translation results to the device.

[0249] 8. The device displays the translation results to the user.

[0250] Specific examples

[0251] If the user types "Hello, what time is it?" and the tone of voice is detected as excited, the following example prompt will be sent by the device:

[0252] Text input: "Hello, what time is it?"

[0253] Audio Tone Information: "Excited"

[0254] Desired language: "English"

[0255] Based on this, the system generates the translation result "What time is it right now?" and provides it to the user.

[0256] This system enables real-time, emotionally-reflective multilingual communication in the business and tourism sectors, and is equipped with highly accurate error handling capabilities to provide highly reliable services.

[0257] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0258] Program processing steps

[0259] Step 1: User Input

[0260] A user accesses the system and inputs text or voice into the device's input field. The input can include text such as "Hello, what time is it?". Once the input is complete, the device records the voice tone and the user's facial expression and sends it to the emotion recognition engine.

[0261] Input: User-typed text or speech. "Hello, what time is it?"

[0262] Output: Input text, audio, emotional information such as tone of voice and facial expressions.

[0263] Specific action: The user speaks into the smartphone's microphone or enters text on the keyboard.

[0264] Step 2: Submitting input data

[0265] The device sends the text, voice data, and emotion information entered by the user to the server, including the user ID and desired language information.

[0266] Input: User input data and emotional information.

[0267] Output: The data sent to the server as an API request.

[0268] Specific operation: The terminal sends data to the server using the HTTP protocol.

[0269] Step 3: Data analysis

[0270] The server receives the data from the device and first analyzes the user's input, using natural language processing technology to extract the meaning and intent of the text.

[0271] Input: Text data and emotional information sent from the device.

[0272] Output: Parsed semantic and intent information.

[0273] Specific operation: The server retrieves user information from the database and passes the data to the analysis engine.

[0274] Step 4: Translation process

[0275] The server translates the analyzed text into a specified language (e.g., English) using a generative AI model (e.g., GPT-3). The server inputs the prompt sentence into the generative AI model and obtains the translation result.

[0276] Input: Parsed text data.

[0277] Output: Translated text data. "Hello, what time is it now?"

[0278] Specific operation: The server inputs the prompt sentence into the generative AI model and generates the translation result.

[0279] Step 5: Emotion Recognition

[0280] The server uses an emotion recognition engine to recognize the user's emotions based on the transmitted voice tone and facial expression data. The emotion recognition engine returns emotional information such as "excitement" based on the voice tone and facial expression.

[0281] Input: speech tone and facial expression data.

[0282] Output: Recognized emotion information. "Excited"

[0283] Specific operation: The server sends data to the emotion recognition API and obtains emotion information.

[0284] Step 6: Adjust the translation results

[0285] The server adjusts the translation result based on the emotion recognition results. For example, if the user is excited, the server changes the expression to something like "What time is it right now?"

[0286] Input: Translated text data, emotion information.

[0287] Output: Adjusted translation data. "What time is it right now?"

[0288] Specific operation: The server applies logic to reflect emotions in the translation results.

[0289] Step 7: Send the translation

[0290] The server transmits the adjusted translation result to the terminal.

[0291] Input: Adjusted translation data. "What time is it right now?"

[0292] Output: Data sent to the device as an API response.

[0293] Specific operation: The server sends data to the device.

[0294] Step 8: Displaying the translation results

[0295] The device displays the translation results received from the server to the user, and in some cases plays them back aloud.

[0296] Input: Translation data sent from the server.

[0297] Output: The translation result that is displayed to the user.

[0298] Specific operation: The device displays the translation results on the screen and, if necessary, plays the results aloud using a speech synthesis engine.

[0299] (Application example 2)

[0300] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0301] When communicating smoothly with multilingual customers in a brick-and-mortar store or other setting, real-time translation and appropriate reflection of emotions are necessary. However, many conventional systems only translate the language and are unable to provide translation results that reflect the user's emotions. This can lead to communication inaccuracies and emotional gaps, resulting in lower satisfaction and misunderstandings. Technology to resolve this issue is needed.

[0302] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0303] In this invention, the server includes means for receiving input from a user, means for analyzing the received input, means for recognizing the user's emotion, means for translating the analyzed input into a specified language, means for adjusting the translation result based on the emotion, and means for providing the translated result to the user.

[0304] This makes it possible to perform multilingual translation in real time, incorporating the user's emotions, and facilitate smooth communication.

[0305] "User" means any person or entity that accesses and provides input to the System.

[0306] "Input" refers to data that a user provides to a system, either textually or vocally.

[0307] "Analysis" refers to the process of examining and analyzing received input data and preparing it for language translation or emotion recognition.

[0308] The "specified language" refers to the language to which the user wishes to translate, and is a specific language selected from among multiple languages.

[0309] "Translation" refers to the process of converting text or audio expressed in one language into a different language.

[0310] "Emotion" indicates the user's psychological state and includes information recognized from non-verbal elements such as tone of voice and facial expression.

[0311] "Adjustment" refers to changing the translated result to an appropriate expression based on the user's feelings.

[0312] "Terminal" refers to a device used by a user to access the system, such as a smartphone or head-mounted display.

[0313] "Providing" refers to the act of displaying or presenting the processed translation results to the user.

[0314] "Real-time" means that the time between input and translation or adjustment results being provided is very short, with an immediate response.

[0315] This invention is a system that translates text and speech input by a user into other languages ​​in real time, and further recognizes the user's emotions and reflects them in the translation results, facilitating smooth communication. This system is mainly composed of four entities: a server, a terminal, a user, and an emotion engine. Specifically, the system operates as follows:

[0316] Hardware and Software Configuration

[0317] User

[0318] Users access the system through their own devices and input text or voice into the input field. When users input voice, it is converted into text data using voice recognition software (e.g., Google API). At the same time, the user's tone of voice and facial expressions are analyzed to generate emotion data.

[0319] Terminal

[0320] The terminal consists of a device such as a smartphone or a head-mounted display (HMD). The terminal receives input from the user, analyzes the data and emotional information, and sends it to the server. The terminal has a speech recognition library (e.g., the speech_recognition library) and video analysis software installed.

[0321] server

[0322] The server is responsible for analyzing the input data it receives, specifically:

[0323] Translates received text data into a specified language using a generative model (e.g., Hugging Face's transformers library).

[0324] An emotion engine is used to recognize the user's emotions and generate emotion data.

[0325] The translation results are adjusted based on the user's emotional information.

[0326] Emotion Engine

[0327] The emotion engine is a program that recognizes emotions by analyzing the user's tone of voice, facial expressions, etc. The emotion engine is configured using, for example, the sentiment_analysis_toolkit library. This allows it to obtain emotional information such as whether the user is excited or calm.

[0328] Processing Flow

[0329] 1. User Input

[0330] The user speaks, "Hello, what time is it?" At the same time, the user's tone of voice is analyzed to detect excitement.

[0331] 2. Device Operation

[0332] The device converts the voice into text and sends the input data and emotional information to the server.

[0333] 3. Server Processing

[0334] The server analyzes the received input and uses a generative model to translate it as "Hello, what time is it now?". At the same time, it adjusts the translation result to "What time is it right now?" based on emotional information (excitement state).

[0335] 4. Providing translation results

[0336] The server sends the translated and adjusted results to the device, which then displays them to the user.

[0337] This allows users to communicate smoothly in real time with people who speak other languages.

[0338] Specific examples

[0339] For example, if a staff member in a physical store says, "Can I help you with anything?", the voice is recognized and translated into English on the server. Then, based on the user's emotional data, the translation result is adjusted to "Is there anything I can assist you with right now?" and displayed on the staff member's device. This results in smooth communication with the customer.

[0340] Prompt Sentence Examples

[0341] Japanese text: "Hello, what time is it?"

[0342] User Sentiment: "Excited"

[0343] Target language: "English"

[0344] Sample output: "What time is it right now?"

[0345] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0346] Step 1:

[0347] The user speaks, "Hello, what time is it now?" The user's device receives this voice data and converts it into text data using speech recognition software (e.g., Google API). The output is the text data "Hello, what time is it now?"

[0348] Step 2:

[0349] The device analyzes the tone of the user's voice from the audio data and generates emotion data using an emotion engine (e.g., the sentiment_analysis_toolkit library). As a result of the analysis, it is detected that the user is excited, and this is output as emotion information. The output is emotion data for "excited state."

[0350] Step 3:

[0351] The terminal transmits the converted text data and emotion data to the server. The input is the text data "Hello, what time is it now?" and the emotion data "excited."

[0352] Step 4:

[0353] The server parses the received text data and translates it into the specified language (in this case, English) using a generative AI model (for example, Hugging Face's transformers library). The input is the text data "Hello, what time is it now?", and the output is the translation result "Hello, what time is it now?".

[0354] Step 5:

[0355] The server uses an emotion engine to analyze the received emotion data and adjust the translation result. The input is the translation text "Hello, what time is it now?" and the emotion data "excited," and the output is the adjusted translation result of "What time is it right now?"

[0356] Step 6:

[0357] The server sends the adjusted translation result to the terminal. The input is the adjusted translation result of "What time is it right now?" and transfers it to the terminal.

[0358] Step 7:

[0359] The device displays the translation results it receives to the user. The input is the adjusted translation result, "What time is it right now?", which is displayed visually to the user. The output is the message, "What time is it right now?", which the user sees on the device.

[0360] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0361] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0362] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0363] [Second embodiment]

[0364] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0365] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0366] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0367] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0368] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0369] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0370] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0371] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0372] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0373] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0374] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0375] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0376] This invention is a system that translates text and speech input by a user into other languages ​​in real time, facilitating communication. This system is mainly composed of three entities: a server, a terminal, and a user.

[0377] overview

[0378] A user accesses the system and enters text or voice into an input field. For example, a user might enter "Hello, what time is it?"

[0379] The device receives the user's input and sends it to the server, which analyzes the received data and translates it into a specified language (e.g., English) using a generative model.

[0380] Once the translation is complete, the server sends the translation results to the terminal, which then displays them to the user, allowing users to communicate smoothly with users of other languages.

[0381] Specific explanation of each element

[0382] User Input

[0383] Users access the platform through their own devices and enter text or voice into the input fields, which can be used in a variety of situations, such as business meetings or exchanging information at tourist spots.

[0384] For example, consider the case where a user uses their smartphone to type "Hello, what time is it now?" in Japanese.

[0385] Device operation

[0386] The device receives the user's input in real time and transmits the data to the server, for example via its own HTTP protocol, including the user's input text, user ID, and desired language information.

[0387] Server Processing

[0388] The server analyzes the received data. Once the analysis is complete, it uses a generative model to perform the translation process. The generative model utilizes the latest AI technology to achieve high accuracy and real-time translation.

[0389] For example, the server parses the Japanese phrase "Hello, what time is it now?" and translates it into English using the appropriate model, resulting in the translation "Hello, what time is it now?"

[0390] Providing translation results

[0391] After the translation is complete, the server sends the results to the terminal, which then displays them to the user in a user-friendly format, making them instantly available.

[0392] Specific examples

[0393] As a concrete example, the translation flow from Japanese to English is shown below.

[0394] The user types, "Hello, what time is it?"

[0395] The terminal sends the input to the server.

[0396] The server parses the received input and translates it into "Hello, what time is it now?"

[0397] The server sends the translation results to the terminal.

[0398] The terminal displays the translation results to the user.

[0399] This allows users to communicate smoothly with speakers of other languages ​​in real time. This system is particularly expected to be used in the business and tourism sectors, and its error handling functionality will provide a highly reliable service.

[0400] The processing flow will be explained below.

[0401] Step 1:

[0402] A user accesses the platform and enters text or voice into an input field, for example, "Hello, what time is it?"

[0403] Step 2:

[0404] The device receives the user's input and sends the data to the server, including the user's input text, user ID, and desired language information.

[0405] Step 3:

[0406] The server then analyzes the received data, which includes identifying the language of the text entered by the user and performing any necessary pre-processing (e.g., noise removal and text normalization).

[0407] Step 4:

[0408] The server uses the generative model to translate the parsed input data into the specified language. For example, to translate the Japanese phrase "Hello, what time is it now?" into English, it would be converted to "Hello, what time is it now?"

[0409] Step 5:

[0410] The server formats the translation results and sends them to the device, ensuring that the formatted data is organized in a way that is easy for the user to understand.

[0411] Step 6:

[0412] The device receives the translation and displays it to the user, properly formatted in a user-friendly language.

[0413] Step 7:

[0414] The user checks the displayed translation result and takes the next action (re-entering or other operation). If there is further input, the cycle starts again from step 1.

[0415] Through this series of processes, users can communicate smoothly and in real time in other languages.

[0416] Example 1

[0417] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0418] Real-time communication between multiple languages ​​is important, especially in business and tourism. However, conventional translation systems often suffer from low translation accuracy and lack real-time capabilities, making them prone to errors. To address this issue, a system that provides high-accuracy, real-time translation and is easy for users to use is needed.

[0419] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0420] In this invention, the server includes means for parsing received input, means for using a generative AI model to translate the parsed input into a specified language, and means for providing a prompt sentence to the generative AI model, thereby enabling highly accurate and real-time translation.

[0421] "Input" is text or voice data that a user provides to the system via a terminal.

[0422] "Analysis" is the process by which the server understands the input data it receives and extracts the necessary information.

[0423] A "generative AI model" is an artificial intelligence model that uses technologies such as neural networks to translate input data into a specified language.

[0424] A "prompt sentence" is a sentence that instructs the generative AI model on what processing to do.

[0425] "Translation" is the process of converting input data into a specified language.

[0426] "Providing" is the process in which the server sends the translation result to the terminal, and the terminal displays the result to the user.

[0427] A "terminal" is a device (e.g., a smartphone or tablet) that a user uses to access the system and input data or display results.

[0428] "Server" means a central processing unit that receives input from a terminal and performs analysis, translation, and provides results.

[0429] "Result display" is the process by which the terminal visually presents the translation results received from the server to the user.

[0430] The present invention provides a system for translating text and speech input by a user into other languages ​​in real time to facilitate communication. The system operates in cooperation with a user, a terminal, and a server. A specific embodiment of the system is described below.

[0431] User Input

[0432] Users access the system using devices such as smartphones or tablets. They enter text into the provided input fields or use a microphone to input voice. For example, consider a case where a user enters "Hello, what time is it now?" in Japanese.

[0433] Device operation

[0434] The device receives user input in real time and sends the data to the server. The device formats the user input data and adds necessary metadata such as the user ID and the language to be translated. This data is sent to the server via the HTTP protocol. For example, the device generates an HTTP request containing the text data "Hello, what time is it now?" and sends it to the server.

[0435] Server Processing

[0436] The server receives the request sent from the device and extracts the input text and necessary metadata from the request body. After verifying that the received data is free of errors, it performs the translation using a generative AI model. Specifically, the server generates a prompt sentence for the generative AI model (e.g., OpenAI's GPT-4) and passes it on. An example of a prompt sentence is, "Please translate from Japanese to English. Input sentence: 'Hello, what time is it now?'"

[0437] Providing translation results

[0438] Once the translation process is complete, the server sends the translation result to the terminal as an HTTP response. The response includes the translated text (e.g., "Hello, what time is it now?"). At this time, the response is also checked for formatting and errors.

[0439] Displaying results on your device

[0440] The device analyzes the response received from the server and displays the results to the user. For example, the translation result "Hello, what time is it now?" is displayed on the device's display. The user can check the result in real time.

[0441] Specific examples

[0442] Below is a concrete Japanese to English translation flow:

[0443] 1. The user types "Hello, what time is it?" into the text field.

[0444] 2. The device sends the input to the server along with the "User ID" and "Translation Language: English."

[0445] 3. The server receives the data "Hello, what time is it now?" and checks the content.

[0446] 4. The server sends the generative AI model a prompt: "Please translate from Japanese to English. Input sentence: 'Hello, what time is it now?'" and receives a response from the generative AI model: "Hello, what time is it now?"

[0447] 5. The server sends the translation result, "Hello, what time is it now?" to the terminal.

[0448] 6. The terminal displays the result to the user: "Hello, what time is it now?"

[0449] This system will enable users to communicate smoothly with speakers of other languages ​​in real time, and is expected to be particularly useful in the fields of business and tourism.

[0450] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0451] Step 1: User Input

[0452] A user accesses the system through his / her own terminal. The user inputs text or voice into an input field. Specifically, if the user wants to input "Hello, what time is it now?" in Japanese, he / she types the characters into the input field of the terminal or captures the voice using the microphone. The input data (text or voice) is prepared.

[0453] Step 2: Send data from the device to the server

[0454] The terminal formats the user's input data and adds the necessary metadata (user ID, language to translate, etc.), then sends the formatted request to the server using the HTTP protocol. The input data and metadata are sent to the server.

[0455] Step 3: Data analysis on the server

[0456] The server receives the request from the terminal and analyzes its contents. Specifically, it extracts the input text, user ID, and desired language information from the request body. The server then verifies that the data format and content are correct and performs error checking. The analysis results in structured data.

[0457] Step 4: Translation on the server

[0458] The server performs translation using a generative AI model based on the analysis results. Specifically, it generates a prompt sentence for the generative AI model (e.g., generative model GPT-4) and passes it on. For example, a prompt sentence such as "Please translate from Japanese to English. Input sentence: 'Hello, what time is it now?'" is generated and input into the AI ​​model. The generative AI model converts the input data into the specified language and outputs the translated data, "Hello, what time is it now?"

[0459] Step 5: Send the translation results from the server to the device

[0460] The server formats the translation results obtained from the generative AI model as an HTTP response and sends it to the device. The response includes the translation result (e.g., "Hello, what time is it now?"). The server checks the response format and content for errors. The formatted response data is sent to the device.

[0461] Step 6: View the results in your device

[0462] The device analyzes the response received from the server and extracts the translation result. Specifically, the device extracts the translated text from the response data and displays it to the user. The displayed result (e.g., "Hello, what time is it now?") is confirmed by the user.

[0463] Through these steps, users can obtain highly accurate translation results into other languages ​​in real time.

[0464] (Application example 1)

[0465] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0466] In modern society, situations requiring multilingual communication in brick-and-mortar stores are on the rise, but language barriers can make smooth customer service difficult. This problem is particularly pronounced in regions where the number of foreign language speakers is increasing, requiring accurate translation in real time. However, existing translation systems often lack sufficient translation speed and accuracy. Furthermore, they lack the ability to check history, making it impossible to refer to past communications, which risks reducing the quality of service.

[0467] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0468] In this invention, the server includes means for receiving input from a user, means for analyzing the received input, means for translating the analyzed input into a specified language, means for providing the translated result to the user, and means for providing an interface for users and foreign language speakers to communicate in real time at the store. This enables accurate multilingual communication in real time. The inclusion of a history function also allows past communications to be referenced, improving the quality of service. Furthermore, translation is performed using a generative model, and accuracy is improved by specific prompt sentences, providing more reliable translation results.

[0469] "User" refers to a person who uses the system to input text or voice and receive the results translated into another language.

[0470] "Input" refers to the text or voice information provided by a user to a system.

[0471] "Analysis" refers to the information processing means that processes received input data and makes sense of it.

[0472] "Translation" refers to the process of converting parsed input data into another specified language.

[0473] "Interface" refers to the operation screen and operation means that facilitate smooth interaction between the user and the system.

[0474] The "history function" refers to a function that saves past communication details with users and allows them to refer to them as needed.

[0475] "Generative model" refers to a machine learning algorithm that uses AI technology to produce highly accurate translations.

[0476] A "prompt sentence" refers to an instruction sentence given to a generative model to perform an appropriate translation.

[0477] To implement this invention, it is necessary to build a system in which users, terminals, and servers work together to provide highly accurate real-time translation and support smooth communication between multiple languages.

[0478] User operations

[0479] Users access the system using devices such as smartphones or tablets. The device displays an interface as an operation screen, where users input text or voice. For example, consider the case where a store clerk asks a foreign customer, "Hello, how do you want to use this product?" in Japanese.

[0480] Device operation

[0481] The terminal receives text and voice input from the user and sends it to the server, generating a data packet containing the user's input data, user ID, and desired language information using a communication method such as the HTTP protocol.

[0482] Server Processing

[0483] The server analyzes the input data sent from the device. After analysis, it uses a generative AI model to translate it into the specified language. Specifically, it uses OpenAI's API to translate based on the following prompt:

[0484] Prompt Sentence Examples

[0485] "Translate this text to English: Hello, how do I use this product?"

[0486] Based on this prompt, the generative model generates a translation result such as "Hello, how do I use this product?"

[0487] Providing translation results

[0488] The server then sends the translation result back to the terminal, which then displays it in a format that is easy for the user to understand. For example, the clerk's terminal screen might display "Hello, how do I use this product?" in English.

[0489] History function

[0490] Furthermore, the system has a built-in history function that allows past communications to be saved, allowing users to refer to past conversation history as needed to improve the quality of service.

[0491] Hardware and Software

[0492] Hardware: Smartphone or tablet (iOS or Android)

[0493] Software: The front end uses HTML, CSS, and JavaScript, and the back end uses Python Flask and the OpenAI API.

[0494] By combining these steps, the system facilitates multilingual communication in brick-and-mortar stores, enabling users and foreign language speakers to communicate effectively in real time.

[0495] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0496] Step 1:

[0497] A user accesses the system using a device such as a smartphone or tablet and inputs text or voice into an input field. For example, a store clerk might input "Hello, how do you want to use this product?" in Japanese. In this case, the input is sent to the device as text or voice data.

[0498] Step 2:

[0499] The device analyzes the received text and voice data and sends it to the server. At this point, a data packet is formed using the HTTP protocol or similar, and information such as the user ID and desired language (e.g., English) is also included. This is the process by which the text and voice data received as input is converted into a data packet to be sent to the server.

[0500] Step 3:

[0501] The server receives data packets sent from the device and analyzes the input data. After analysis, it uses a generative AI model to translate it into the specified language. For example, the server provides the prompt sentence "Translate this text to English: Hello, how do you use this product?" to the generative model, and the resulting translation is "Hello, how do I use this product?"

[0502] Step 4:

[0503] The server then sends the translation results back to the terminal. This time, the translation results are sent in text format. The server's input is the translated text data, and it is the process that processes it to send it to the terminal.

[0504] Step 5:

[0505] The terminal receives the translation results from the server and displays them to the user. The results are displayed visually through a user-friendly interface, allowing the user to understand them immediately. For example, the salesperson's terminal screen might display "Hello, how do I use this product?"

[0506] Step 6:

[0507] As a history function, the device will store past translation history and allow users to refer to this history as needed, supporting continuous communication with the same user and improving the quality of service. The history data will be stored in local storage or cloud storage, allowing users to access it at any time.

[0508] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0509] This invention is a system that translates user-entered text and speech into other languages ​​in real time, and also recognizes the user's emotions and reflects them in the translation results, facilitating smooth communication. This system is mainly composed of four main components: a server, a terminal, a user, and an emotion engine.

[0510] overview

[0511] A user accesses the system and inputs text or voice into an input field. For example, consider the case where a user inputs "Hello, what time is it now?". At the same time, the system reads the user's emotions from their tone of voice and facial expressions.

[0512] The device receives user input and sends the data and emotion information to the server, which analyzes the received data and translates it into a specified language (e.g., English) using a generative model. It also uses an emotion engine to recognize the user's emotion and adjust the translation results based on that emotion.

[0513] Once the translation is complete, the server sends the translation results to the terminal, which then displays them to the user, allowing users to communicate smoothly with users of other languages.

[0514] Specific explanation of each element

[0515] User Input

[0516] Users access the platform through their own devices and enter text or voice into the input field. This is used in a variety of situations, such as business meetings and exchanging information at tourist spots. For example, consider a case where a user uses their smartphone to enter "Hello, what time is it now?" in Japanese.

[0517] Device operation

[0518] The device receives user input in real time and transmits the data and emotion information to the server via its own HTTP protocol, etc. The transmitted data includes the user's input text, user ID, desired language information, and emotion information.

[0519] Server Processing

[0520] The server analyzes the received data. Once the analysis is complete, it uses a generative model to perform the translation process. The generative model utilizes the latest AI technology to achieve high accuracy and real-time translation.

[0521] For example, the server analyzes the Japanese phrase "Hello, what time is it now?" and translates it into English using an appropriate model, resulting in the translation "Hello, what time is it now?"

[0522] In parallel, the server uses an emotion engine to recognize the user's emotions and adjusts the translated text accordingly. For example, if the user is excited, the server might translate it as "What time is it right now?"

[0523] Providing translation results

[0524] After the translation is complete, the server sends the results to the device, which then displays them to the user in a user-friendly format, taking into account emotional information.

[0525] Specific examples

[0526] As a concrete example, the translation flow from Japanese to English is shown below.

[0527] A user types, "Hello, what time is it?" At the same time, the user's tone of voice is analyzed and it is detected that the user is excited.

[0528] The device sends the input and emotion information to the server.

[0529] The server analyzes the received input and translates it as "Hello, what time is it now?". It also adjusts the translation result based on the emotion information to "What time is it right now?"

[0530] The server sends the translation results to the terminal.

[0531] The terminal displays the translation results to the user.

[0532] This allows users to communicate smoothly and emotionally with speakers of other languages ​​in real time. This system is particularly expected to be used in the business and tourism sectors, and its error handling functionality will provide a highly reliable service.

[0533] The processing flow will be explained below.

[0534] Step 1:

[0535] A user accesses the platform and enters text or voice into an input field, for example, "Hello, what time is it?" The user's facial expressions and tone of voice are also recorded.

[0536] Step 2:

[0537] The device receives user input and converts it into text data, while also collecting the user's emotional information (e.g., excitement, sadness, joy).

[0538] Step 3:

[0539] The device sends the user's input data and emotion information to the server, along with the user ID and desired language.

[0540] Step 4:

[0541] The server analyzes the received data, which includes identifying the language of the text and pre-processing it (denoising, normalization).

[0542] Step 5:

[0543] The server uses the generative model to translate the parsed text into the specified language. For example, the Japanese phrase "Hello, what time is it now?" is translated into English as "Hello, what time is it now?"

[0544] Step 6:

[0545] The server uses the emotion engine to recognize the user's emotion. For example, if the user is excited, the emotion is recognized as "excited."

[0546] Step 7:

[0547] The server adjusts the translation results based on the emotional information it recognizes. For example, if the user is feeling excited, it will modify "Hello, what time is it now?" to reflect that emotion, such as "What time is it right now?"

[0548] Step 8:

[0549] The server formats the final translation and sends it to the device, with adjustments based on emotional information.

[0550] Step 9:

[0551] The device will then display the translation results to the user, specifically asking, "What time is it right now?"

[0552] Step 10:

[0553] The user checks the displayed translation result and takes the next action (re-entering or other operations) if necessary. If re-entering is required, the cycle starts again from step 1.

[0554] Through this series of processes, users can communicate smoothly and emotionally with speakers of other languages ​​in real time.

[0555] Example 2

[0556] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0557] In modern society, real-time communication between people who speak different languages ​​is becoming increasingly important. However, conventional translation systems simply translate text and are unable to provide translation results that reflect the user's emotions. This makes it difficult to accurately convey important emotional information in communication.

[0558] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving input from a user, means for analyzing the received input, means for translating the analyzed input into a specified language, means for recognizing the user's emotion, means for adjusting the translation result based on the emotion recognition result, and means for providing the translation result to the user. This makes it possible to provide a translation result that reflects emotional information contained in the text or voice input by the user.

[0559] "User" means any individual or entity that uses the System.

[0560] "Terminal" refers to an electronic device that allows a user to access and operate the system. Examples include smartphones and personal computers.

[0561] A "server" is a computer system that receives data sent from a terminal and analyzes and processes the data.

[0562] "Means for receiving input" refers to a mechanism by which the terminal can receive text or voice data entered by a user.

[0563] "Means for analyzing input" refers to a mechanism for analyzing received data and understanding its meaning and intent. This includes natural language processing technology.

[0564] "Means for translating" means a mechanism for translating analyzed text into another specified language, including a generative AI model.

[0565] "Means for recognizing emotions" refers to a mechanism for detecting emotions from the user's tone of voice and facial expressions and acquiring that information.

[0566] The "means for adjusting the translation result" refers to a mechanism for correcting the translated text based on the recognized emotional information and adding appropriate emotional expressions.

[0567] "Means for providing the translated results to the user" refers to a mechanism for displaying or audibly notifying the user of the final translation results via a terminal.

[0568] A "generative AI model" is an algorithm that uses artificial intelligence techniques to generate or process data. For example, it includes deep learning models that specialize in natural language processing.

[0569] A "prompt sentence" refers to the text data input into a generative AI model, and is the initial information that the model uses to translate and generate.

[0570] This system translates user-input text and speech into other languages ​​in real time, and also recognizes the user's emotions and reflects them in the translation results. This system is mainly composed of four main components: a server, a terminal, a user, and an emotion recognition engine.

[0571] System configuration and technologies used

[0572] User

[0573] A user accesses their device and inputs text or voice into an input field. For example, a user might input "Hello, what time is it now?" in Japanese. The device then sends the user's tone of voice and facial expressions to an emotion recognition engine.

[0574] Terminal

[0575] The device receives user input in real time and transmits the data and emotion information to the server. The device can be a smartphone or a PC, and data transmission uses the HTTP protocol. The transmitted data includes the input text, user ID, desired language information, and emotion information.

[0576] server

[0577] The server analyzes the data received from the device and translates it into the specified language using a generative AI model. Generative AI models use algorithms specialized for natural language processing, such as OpenAI's GPT-3 model. The server uses an emotion recognition engine (such as Microsoft Azure Cognitive Services) to recognize the user's emotions and adjust the translation results based on that information.

[0578] Emotion Recognition Engine

[0579] The emotion recognition engine analyzes the user's tone of voice and facial expression data to obtain emotional information. Based on this emotional information, the server adjusts the translation results and adds appropriate emotional expressions.

[0580] Specific operation flow

[0581] 1. The user types "Hello, what time is it?". At the same time, the user's excitement level is collected as emotional information.

[0582] 2. The device sends the input and emotion information to the server.

[0583] 3. The server receives the request and parses the data.

[0584] 4. The server uses the generative AI model to translate it as "Hello, what time is it now?"

[0585] 5. The server sends the data to the emotion recognition engine to check the user's excitement level.

[0586] 6. The server adjusts the translation result to "What time is it right now?"

[0587] 7. The server sends the translation results to the device.

[0588] 8. The device displays the translation results to the user.

[0589] Specific examples

[0590] If the user types "Hello, what time is it?" and the tone of voice is detected as excited, the following example prompt will be sent by the device:

[0591] Text input: "Hello, what time is it?"

[0592] Audio Tone Information: "Excited"

[0593] Desired language: "English"

[0594] Based on this, the system generates the translation result "What time is it right now?" and provides it to the user.

[0595] This system enables real-time, emotionally-reflective multilingual communication in the business and tourism sectors, and is equipped with highly accurate error handling capabilities to provide highly reliable services.

[0596] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0597] Program processing steps

[0598] Step 1: User Input

[0599] A user accesses the system and inputs text or voice into the device's input field. The input can include text such as "Hello, what time is it?". Once the input is complete, the device records the voice tone and the user's facial expression and sends it to the emotion recognition engine.

[0600] Input: User-typed text or speech. "Hello, what time is it?"

[0601] Output: Input text, audio, emotional information such as tone of voice and facial expressions.

[0602] Specific action: The user speaks into the smartphone's microphone or enters text on the keyboard.

[0603] Step 2: Submitting input data

[0604] The device sends the text, voice data, and emotion information entered by the user to the server, including the user ID and desired language information.

[0605] Input: User input data and emotional information.

[0606] Output: The data sent to the server as an API request.

[0607] Specific operation: The terminal sends data to the server using the HTTP protocol.

[0608] Step 3: Data analysis

[0609] The server receives the data from the device and first analyzes the user's input, using natural language processing technology to extract the meaning and intent of the text.

[0610] Input: Text data and emotional information sent from the device.

[0611] Output: Parsed semantic and intent information.

[0612] Specific operation: The server retrieves user information from the database and passes the data to the analysis engine.

[0613] Step 4: Translation process

[0614] The server translates the analyzed text into a specified language (e.g., English) using a generative AI model (e.g., GPT-3). The server inputs the prompt sentence into the generative AI model and obtains the translation result.

[0615] Input: Parsed text data.

[0616] Output: Translated text data. "Hello, what time is it now?"

[0617] Specific operation: The server inputs the prompt sentence into the generative AI model and generates the translation result.

[0618] Step 5: Emotion Recognition

[0619] The server uses an emotion recognition engine to recognize the user's emotions based on the transmitted voice tone and facial expression data. The emotion recognition engine returns emotional information such as "excitement" based on the voice tone and facial expression.

[0620] Input: speech tone and facial expression data.

[0621] Output: Recognized emotion information. "Excited"

[0622] Specific operation: The server sends data to the emotion recognition API and obtains emotion information.

[0623] Step 6: Adjust the translation results

[0624] The server adjusts the translation result based on the emotion recognition results. For example, if the user is excited, the server changes the expression to something like "What time is it right now?"

[0625] Input: Translated text data, emotion information.

[0626] Output: Adjusted translation data. "What time is it right now?"

[0627] Specific operation: The server applies logic to reflect emotions in the translation results.

[0628] Step 7: Send the translation

[0629] The server transmits the adjusted translation result to the terminal.

[0630] Input: Adjusted translation data. "What time is it right now?"

[0631] Output: Data sent to the device as an API response.

[0632] Specific operation: The server sends data to the device.

[0633] Step 8: Displaying the translation results

[0634] The device displays the translation results received from the server to the user, and in some cases plays them back aloud.

[0635] Input: Translation data sent from the server.

[0636] Output: The translation result that is displayed to the user.

[0637] Specific operation: The device displays the translation results on the screen and, if necessary, plays the results aloud using a speech synthesis engine.

[0638] (Application example 2)

[0639] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0640] When communicating smoothly with multilingual customers in a brick-and-mortar store or other setting, real-time translation and appropriate reflection of emotions are necessary. However, many conventional systems only translate the language and are unable to provide translation results that reflect the user's emotions. This can lead to communication inaccuracies and emotional gaps, resulting in lower satisfaction and misunderstandings. Technology to resolve this issue is needed.

[0641] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0642] In this invention, the server includes means for receiving input from a user, means for analyzing the received input, means for recognizing the user's emotion, means for translating the analyzed input into a specified language, means for adjusting the translation result based on the emotion, and means for providing the translated result to the user.

[0643] This makes it possible to perform multilingual translation in real time, incorporating the user's emotions, and facilitate smooth communication.

[0644] "User" means any person or entity that accesses and provides input to the System.

[0645] "Input" refers to data that a user provides to a system, either textually or vocally.

[0646] "Analysis" refers to the process of examining and analyzing received input data and preparing it for language translation or emotion recognition.

[0647] The "specified language" refers to the language to which the user wishes to translate, and is a specific language selected from among multiple languages.

[0648] "Translation" refers to the process of converting text or audio expressed in one language into a different language.

[0649] "Emotion" indicates the user's psychological state and includes information recognized from non-verbal elements such as tone of voice and facial expression.

[0650] "Adjustment" refers to changing the translated result to an appropriate expression based on the user's feelings.

[0651] "Terminal" refers to a device used by a user to access the system, such as a smartphone or head-mounted display.

[0652] "Providing" refers to the act of displaying or presenting the processed translation results to the user.

[0653] "Real-time" means that the time between input and translation or adjustment results being provided is very short, with an immediate response.

[0654] This invention is a system that translates text and speech input by a user into other languages ​​in real time, and further recognizes the user's emotions and reflects them in the translation results, facilitating smooth communication. This system is mainly composed of four entities: a server, a terminal, a user, and an emotion engine. Specifically, the system operates as follows:

[0655] Hardware and Software Configuration

[0656] User

[0657] Users access the system through their own devices and input text or voice into the input field. When users input voice, it is converted into text data using voice recognition software (e.g., Google API). At the same time, the user's tone of voice and facial expressions are analyzed to generate emotion data.

[0658] Terminal

[0659] The terminal consists of a device such as a smartphone or a head-mounted display (HMD). The terminal receives input from the user, analyzes the data and emotional information, and sends it to the server. The terminal has a speech recognition library (e.g., the speech_recognition library) and video analysis software installed.

[0660] server

[0661] The server is responsible for analyzing the input data it receives, specifically:

[0662] Translates received text data into a specified language using a generative model (e.g., Hugging Face's transformers library).

[0663] An emotion engine is used to recognize the user's emotions and generate emotion data.

[0664] The translation results are adjusted based on the user's emotional information.

[0665] Emotion Engine

[0666] The emotion engine is a program that recognizes emotions by analyzing the user's tone of voice, facial expressions, etc. The emotion engine is configured using, for example, the sentiment_analysis_toolkit library. This allows it to obtain emotional information such as whether the user is excited or calm.

[0667] Processing Flow

[0668] 1. User Input

[0669] The user speaks, "Hello, what time is it?" At the same time, the user's tone of voice is analyzed to detect excitement.

[0670] 2. Device Operation

[0671] The device converts the voice into text and sends the input data and emotional information to the server.

[0672] 3. Server Processing

[0673] The server analyzes the received input and uses a generative model to translate it as "Hello, what time is it now?". At the same time, it adjusts the translation result to "What time is it right now?" based on emotional information (excitement state).

[0674] 4. Providing translation results

[0675] The server sends the translated and adjusted results to the device, which then displays them to the user.

[0676] This allows users to communicate smoothly in real time with people who speak other languages.

[0677] Specific examples

[0678] For example, if a staff member in a physical store says, "Can I help you with anything?", the voice is recognized and translated into English on the server. Then, based on the user's emotional data, the translation result is adjusted to "Is there anything I can assist you with right now?" and displayed on the staff member's device. This results in smooth communication with the customer.

[0679] Prompt Sentence Examples

[0680] Japanese text: "Hello, what time is it?"

[0681] User Sentiment: "Excited"

[0682] Target language: "English"

[0683] Sample output: "What time is it right now?"

[0684] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0685] Step 1:

[0686] The user speaks, "Hello, what time is it now?" The user's device receives this voice data and converts it into text data using speech recognition software (e.g., Google API). The output is the text data "Hello, what time is it now?"

[0687] Step 2:

[0688] The device analyzes the tone of the user's voice from the audio data and generates emotion data using an emotion engine (e.g., the sentiment_analysis_toolkit library). As a result of the analysis, it is detected that the user is excited, and this is output as emotion information. The output is emotion data for "excited state."

[0689] Step 3:

[0690] The terminal transmits the converted text data and emotion data to the server. The input is the text data "Hello, what time is it now?" and the emotion data "excited."

[0691] Step 4:

[0692] The server parses the received text data and translates it into the specified language (in this case, English) using a generative AI model (for example, Hugging Face's transformers library). The input is the text data "Hello, what time is it now?", and the output is the translation result "Hello, what time is it now?".

[0693] Step 5:

[0694] The server uses an emotion engine to analyze the received emotion data and adjust the translation result. The input is the translation text "Hello, what time is it now?" and the emotion data "excited," and the output is the adjusted translation result of "What time is it right now?"

[0695] Step 6:

[0696] The server sends the adjusted translation result to the terminal. The input is the adjusted translation result of "What time is it right now?" and transfers it to the terminal.

[0697] Step 7:

[0698] The device displays the translation results it receives to the user. The input is the adjusted translation result, "What time is it right now?", which is displayed visually to the user. The output is the message, "What time is it right now?", which the user sees on the device.

[0699] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0700] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0701] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0702] [Third embodiment]

[0703] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0704] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0705] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0706] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0707] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0708] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0709] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0710] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0711] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0712] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0713] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0714] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0715] This invention is a system that translates text and speech input by a user into other languages ​​in real time, facilitating communication. This system is mainly composed of three entities: a server, a terminal, and a user.

[0716] overview

[0717] A user accesses the system and enters text or voice into an input field. For example, a user might enter "Hello, what time is it?"

[0718] The device receives the user's input and sends it to the server, which analyzes the received data and translates it into a specified language (e.g., English) using a generative model.

[0719] Once the translation is complete, the server sends the translation results to the terminal, which then displays them to the user, allowing users to communicate smoothly with users of other languages.

[0720] Specific explanation of each element

[0721] User Input

[0722] Users access the platform through their own devices and enter text or voice into the input fields, which can be used in a variety of situations, such as business meetings or exchanging information at tourist spots.

[0723] For example, consider the case where a user uses their smartphone to type "Hello, what time is it now?" in Japanese.

[0724] Device operation

[0725] The device receives the user's input in real time and transmits the data to the server, for example via its own HTTP protocol, including the user's input text, user ID, and desired language information.

[0726] Server Processing

[0727] The server analyzes the received data. Once the analysis is complete, it uses a generative model to perform the translation process. The generative model utilizes the latest AI technology to achieve high accuracy and real-time translation.

[0728] For example, the server parses the Japanese phrase "Hello, what time is it now?" and translates it into English using the appropriate model, resulting in the translation "Hello, what time is it now?"

[0729] Providing translation results

[0730] After the translation is complete, the server sends the results to the terminal, which then displays them to the user in a user-friendly format, making them instantly available.

[0731] Specific examples

[0732] As a concrete example, the translation flow from Japanese to English is shown below.

[0733] The user types, "Hello, what time is it?"

[0734] The terminal sends the input to the server.

[0735] The server parses the received input and translates it into "Hello, what time is it now?"

[0736] The server sends the translation results to the terminal.

[0737] The terminal displays the translation results to the user.

[0738] This allows users to communicate smoothly with speakers of other languages ​​in real time. This system is particularly expected to be used in the business and tourism sectors, and its error handling functionality will provide a highly reliable service.

[0739] The processing flow will be explained below.

[0740] Step 1:

[0741] A user accesses the platform and enters text or voice into an input field, for example, "Hello, what time is it?"

[0742] Step 2:

[0743] The device receives the user's input and sends the data to the server, including the user's input text, user ID, and desired language information.

[0744] Step 3:

[0745] The server then analyzes the received data, which includes identifying the language of the text entered by the user and performing any necessary pre-processing (e.g., noise removal and text normalization).

[0746] Step 4:

[0747] The server uses the generative model to translate the parsed input data into the specified language. For example, to translate the Japanese phrase "Hello, what time is it now?" into English, it would be converted to "Hello, what time is it now?"

[0748] Step 5:

[0749] The server formats the translation results and sends them to the device, ensuring that the formatted data is organized in a way that is easy for the user to understand.

[0750] Step 6:

[0751] The device receives the translation and displays it to the user, properly formatted in a user-friendly language.

[0752] Step 7:

[0753] The user checks the displayed translation result and takes the next action (re-entering or other operation). If there is further input, the cycle starts again from step 1.

[0754] Through this series of processes, users can communicate smoothly and in real time in other languages.

[0755] Example 1

[0756] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0757] Real-time communication between multiple languages ​​is important, especially in business and tourism. However, conventional translation systems often suffer from low translation accuracy and lack real-time capabilities, making them prone to errors. To address this issue, a system that provides high-accuracy, real-time translation and is easy for users to use is needed.

[0758] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0759] In this invention, the server includes means for parsing received input, means for using a generative AI model to translate the parsed input into a specified language, and means for providing a prompt sentence to the generative AI model, thereby enabling highly accurate and real-time translation.

[0760] "Input" is text or voice data that a user provides to the system via a terminal.

[0761] "Analysis" is the process by which the server understands the input data it receives and extracts the necessary information.

[0762] A "generative AI model" is an artificial intelligence model that uses technologies such as neural networks to translate input data into a specified language.

[0763] A "prompt sentence" is a sentence that instructs the generative AI model on what processing to do.

[0764] "Translation" is the process of converting input data into a specified language.

[0765] "Providing" is the process in which the server sends the translation result to the terminal, and the terminal displays the result to the user.

[0766] A "terminal" is a device (e.g., a smartphone or tablet) that a user uses to access the system and input data or display results.

[0767] "Server" means a central processing unit that receives input from a terminal and performs analysis, translation, and provides results.

[0768] "Result display" is the process by which the terminal visually presents the translation results received from the server to the user.

[0769] The present invention provides a system for translating text and speech input by a user into other languages ​​in real time to facilitate communication. The system operates in cooperation with a user, a terminal, and a server. A specific embodiment of the system is described below.

[0770] User Input

[0771] Users access the system using devices such as smartphones or tablets. They enter text into the provided input fields or use a microphone to input voice. For example, consider a case where a user enters "Hello, what time is it now?" in Japanese.

[0772] Device operation

[0773] The device receives user input in real time and sends the data to the server. The device formats the user input data and adds necessary metadata such as the user ID and the language to be translated. This data is sent to the server via the HTTP protocol. For example, the device generates an HTTP request containing the text data "Hello, what time is it now?" and sends it to the server.

[0774] Server Processing

[0775] The server receives the request sent from the device and extracts the input text and necessary metadata from the request body. After verifying that the received data is free of errors, it performs the translation using a generative AI model. Specifically, the server generates a prompt sentence for the generative AI model (e.g., OpenAI's GPT-4) and passes it on. An example of a prompt sentence is, "Please translate from Japanese to English. Input sentence: 'Hello, what time is it now?'"

[0776] Providing translation results

[0777] Once the translation process is complete, the server sends the translation result to the terminal as an HTTP response. The response includes the translated text (e.g., "Hello, what time is it now?"). At this time, the response is also checked for formatting and errors.

[0778] Displaying results on your device

[0779] The device analyzes the response received from the server and displays the results to the user. For example, the translation result "Hello, what time is it now?" is displayed on the device's display. The user can check the result in real time.

[0780] Specific examples

[0781] Below is a concrete Japanese to English translation flow:

[0782] 1. The user types "Hello, what time is it?" into the text field.

[0783] 2. The device sends the input to the server along with the "User ID" and "Translation Language: English."

[0784] 3. The server receives the data "Hello, what time is it now?" and checks the content.

[0785] 4. The server sends the generative AI model a prompt: "Please translate from Japanese to English. Input sentence: 'Hello, what time is it now?'" and receives a response from the generative AI model: "Hello, what time is it now?"

[0786] 5. The server sends the translation result, "Hello, what time is it now?" to the terminal.

[0787] 6. The terminal displays the result to the user: "Hello, what time is it now?"

[0788] This system will enable users to communicate smoothly with speakers of other languages ​​in real time, and is expected to be particularly useful in the fields of business and tourism.

[0789] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0790] Step 1: User Input

[0791] A user accesses the system through his / her own terminal. The user inputs text or voice into an input field. Specifically, if the user wants to input "Hello, what time is it now?" in Japanese, he / she types the characters into the input field of the terminal or captures the voice using the microphone. The input data (text or voice) is prepared.

[0792] Step 2: Send data from the device to the server

[0793] The terminal formats the user's input data and adds the necessary metadata (user ID, language to translate, etc.), then sends the formatted request to the server using the HTTP protocol. The input data and metadata are sent to the server.

[0794] Step 3: Data analysis on the server

[0795] The server receives the request from the terminal and analyzes its contents. Specifically, it extracts the input text, user ID, and desired language information from the request body. The server then verifies that the data format and content are correct and performs error checking. The analysis results in structured data.

[0796] Step 4: Translation on the server

[0797] The server performs translation using a generative AI model based on the analysis results. Specifically, it generates a prompt sentence for the generative AI model (e.g., generative model GPT-4) and passes it on. For example, a prompt sentence such as "Please translate from Japanese to English. Input sentence: 'Hello, what time is it now?'" is generated and input into the AI ​​model. The generative AI model converts the input data into the specified language and outputs the translated data, "Hello, what time is it now?"

[0798] Step 5: Send the translation results from the server to the device

[0799] The server formats the translation results obtained from the generative AI model as an HTTP response and sends it to the device. The response includes the translation result (e.g., "Hello, what time is it now?"). The server checks the response format and content for errors. The formatted response data is sent to the device.

[0800] Step 6: View the results in your device

[0801] The device analyzes the response received from the server and extracts the translation result. Specifically, the device extracts the translated text from the response data and displays it to the user. The displayed result (e.g., "Hello, what time is it now?") is confirmed by the user.

[0802] Through these steps, users can obtain highly accurate translation results into other languages ​​in real time.

[0803] (Application example 1)

[0804] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0805] In modern society, situations requiring multilingual communication in brick-and-mortar stores are on the rise, but language barriers can make smooth customer service difficult. This problem is particularly pronounced in regions where the number of foreign language speakers is increasing, requiring accurate translation in real time. However, existing translation systems often lack sufficient translation speed and accuracy. Furthermore, they lack the ability to check history, making it impossible to refer to past communications, which risks reducing the quality of service.

[0806] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0807] In this invention, the server includes means for receiving input from a user, means for analyzing the received input, means for translating the analyzed input into a specified language, means for providing the translated result to the user, and means for providing an interface for users and foreign language speakers to communicate in real time at the store. This enables accurate multilingual communication in real time. The inclusion of a history function also allows past communications to be referenced, improving the quality of service. Furthermore, translation is performed using a generative model, and accuracy is improved by specific prompt sentences, providing more reliable translation results.

[0808] "User" refers to a person who uses the system to input text or voice and receive the results translated into another language.

[0809] "Input" refers to the text or voice information provided by a user to a system.

[0810] "Analysis" refers to the information processing means that processes received input data and makes sense of it.

[0811] "Translation" refers to the process of converting parsed input data into another specified language.

[0812] "Interface" refers to the operation screen and operation means that facilitate smooth interaction between the user and the system.

[0813] The "history function" refers to a function that saves past communication details with users and allows them to refer to them as needed.

[0814] "Generative model" refers to a machine learning algorithm that uses AI technology to produce highly accurate translations.

[0815] A "prompt sentence" refers to an instruction sentence given to a generative model to perform an appropriate translation.

[0816] To implement this invention, it is necessary to build a system in which users, terminals, and servers work together to provide highly accurate real-time translation and support smooth communication between multiple languages.

[0817] User operations

[0818] Users access the system using devices such as smartphones or tablets. The device displays an interface as an operation screen, where users input text or voice. For example, consider the case where a store clerk asks a foreign customer, "Hello, how do you want to use this product?" in Japanese.

[0819] Device operation

[0820] The terminal receives text and voice input from the user and sends it to the server, generating a data packet containing the user's input data, user ID, and desired language information using a communication method such as the HTTP protocol.

[0821] Server Processing

[0822] The server analyzes the input data sent from the device. After analysis, it uses a generative AI model to translate it into the specified language. Specifically, it uses OpenAI's API to translate based on the following prompt:

[0823] Prompt Sentence Examples

[0824] "Translate this text to English: Hello, how do I use this product?"

[0825] Based on this prompt, the generative model generates a translation result such as "Hello, how do I use this product?"

[0826] Providing translation results

[0827] The server then sends the translation result back to the terminal, which then displays it in a format that is easy for the user to understand. For example, the clerk's terminal screen might display "Hello, how do I use this product?" in English.

[0828] History function

[0829] Furthermore, the system has a built-in history function that allows past communications to be saved, allowing users to refer to past conversation history as needed to improve the quality of service.

[0830] Hardware and Software

[0831] Hardware: Smartphone or tablet (iOS or Android)

[0832] Software: The front end uses HTML, CSS, and JavaScript, and the back end uses Python Flask and the OpenAI API.

[0833] By combining these steps, the system facilitates multilingual communication in brick-and-mortar stores, enabling users and foreign language speakers to communicate effectively in real time.

[0834] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0835] Step 1:

[0836] A user accesses the system using a device such as a smartphone or tablet and inputs text or voice into an input field. For example, a store clerk might input "Hello, how do you want to use this product?" in Japanese. In this case, the input is sent to the device as text or voice data.

[0837] Step 2:

[0838] The device analyzes the received text and voice data and sends it to the server. At this point, a data packet is formed using the HTTP protocol or similar, and information such as the user ID and desired language (e.g., English) is also included. This is the process by which the text and voice data received as input is converted into a data packet to be sent to the server.

[0839] Step 3:

[0840] The server receives data packets sent from the device and analyzes the input data. After analysis, it uses a generative AI model to translate it into the specified language. For example, the server provides the prompt sentence "Translate this text to English: Hello, how do you use this product?" to the generative model, and the resulting translation is "Hello, how do I use this product?"

[0841] Step 4:

[0842] The server then sends the translation results back to the terminal. This time, the translation results are sent in text format. The server's input is the translated text data, and it is the process that processes it to send it to the terminal.

[0843] Step 5:

[0844] The terminal receives the translation results from the server and displays them to the user. The results are displayed visually through a user-friendly interface, allowing the user to understand them immediately. For example, the salesperson's terminal screen might display "Hello, how do I use this product?"

[0845] Step 6:

[0846] As a history function, the device will store past translation history and allow users to refer to this history as needed, supporting continuous communication with the same user and improving the quality of service. The history data will be stored in local storage or cloud storage, allowing users to access it at any time.

[0847] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0848] This invention is a system that translates user-entered text and speech into other languages ​​in real time, and also recognizes the user's emotions and reflects them in the translation results, facilitating smooth communication. This system is mainly composed of four main components: a server, a terminal, a user, and an emotion engine.

[0849] overview

[0850] A user accesses the system and inputs text or voice into an input field. For example, consider the case where a user inputs "Hello, what time is it now?". At the same time, the system reads the user's emotions from their tone of voice and facial expressions.

[0851] The device receives user input and sends the data and emotion information to the server, which analyzes the received data and translates it into a specified language (e.g., English) using a generative model. It also uses an emotion engine to recognize the user's emotion and adjust the translation results based on that emotion.

[0852] Once the translation is complete, the server sends the translation results to the terminal, which then displays them to the user, allowing users to communicate smoothly with users of other languages.

[0853] Specific explanation of each element

[0854] User Input

[0855] Users access the platform through their own devices and enter text or voice into the input field. This is used in a variety of situations, such as business meetings and exchanging information at tourist spots. For example, consider a case where a user uses their smartphone to enter "Hello, what time is it now?" in Japanese.

[0856] Device operation

[0857] The device receives user input in real time and transmits the data and emotion information to the server via its own HTTP protocol, etc. The transmitted data includes the user's input text, user ID, desired language information, and emotion information.

[0858] Server Processing

[0859] The server analyzes the received data. Once the analysis is complete, it uses a generative model to perform the translation process. The generative model utilizes the latest AI technology to achieve high accuracy and real-time translation.

[0860] For example, the server analyzes the Japanese phrase "Hello, what time is it now?" and translates it into English using an appropriate model, resulting in the translation "Hello, what time is it now?"

[0861] In parallel, the server uses an emotion engine to recognize the user's emotions and adjusts the translated text accordingly. For example, if the user is excited, the server might translate it as "What time is it right now?"

[0862] Providing translation results

[0863] After the translation is complete, the server sends the results to the device, which then displays them to the user in a user-friendly format, taking into account emotional information.

[0864] Specific examples

[0865] As a concrete example, the translation flow from Japanese to English is shown below.

[0866] A user types, "Hello, what time is it?" At the same time, the user's tone of voice is analyzed and it is detected that the user is excited.

[0867] The device sends the input and emotion information to the server.

[0868] The server analyzes the received input and translates it as "Hello, what time is it now?". It also adjusts the translation result based on the emotion information to "What time is it right now?"

[0869] The server sends the translation results to the terminal.

[0870] The terminal displays the translation results to the user.

[0871] This allows users to communicate smoothly and emotionally with speakers of other languages ​​in real time. This system is particularly expected to be used in the business and tourism sectors, and its error handling functionality will provide a highly reliable service.

[0872] The processing flow will be explained below.

[0873] Step 1:

[0874] A user accesses the platform and enters text or voice into an input field, for example, "Hello, what time is it?" The user's facial expressions and tone of voice are also recorded.

[0875] Step 2:

[0876] The device receives user input and converts it into text data, while also collecting the user's emotional information (e.g., excitement, sadness, joy).

[0877] Step 3:

[0878] The device sends the user's input data and emotion information to the server, along with the user ID and desired language.

[0879] Step 4:

[0880] The server analyzes the received data, which includes identifying the language of the text and pre-processing it (denoising, normalization).

[0881] Step 5:

[0882] The server uses the generative model to translate the parsed text into the specified language. For example, the Japanese phrase "Hello, what time is it now?" is translated into English as "Hello, what time is it now?"

[0883] Step 6:

[0884] The server uses the emotion engine to recognize the user's emotion. For example, if the user is excited, the emotion is recognized as "excited."

[0885] Step 7:

[0886] The server adjusts the translation results based on the emotional information it recognizes. For example, if the user is feeling excited, it will modify "Hello, what time is it now?" to reflect that emotion, such as "What time is it right now?"

[0887] Step 8:

[0888] The server formats the final translation and sends it to the device, with adjustments based on emotional information.

[0889] Step 9:

[0890] The device will then display the translation results to the user, specifically asking, "What time is it right now?"

[0891] Step 10:

[0892] The user checks the displayed translation result and takes the next action (re-entering or other operations) if necessary. If re-entering is required, the cycle starts again from step 1.

[0893] Through this series of processes, users can communicate smoothly and emotionally with speakers of other languages ​​in real time.

[0894] Example 2

[0895] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0896] In modern society, real-time communication between people who speak different languages ​​is becoming increasingly important. However, conventional translation systems simply translate text and are unable to provide translation results that reflect the user's emotions. This makes it difficult to accurately convey important emotional information in communication.

[0897] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving input from a user, means for analyzing the received input, means for translating the analyzed input into a specified language, means for recognizing the user's emotion, means for adjusting the translation result based on the emotion recognition result, and means for providing the translation result to the user. This makes it possible to provide a translation result that reflects emotional information contained in the text or voice input by the user.

[0898] "User" means any individual or entity that uses the System.

[0899] "Terminal" refers to an electronic device that allows a user to access and operate the system. Examples include smartphones and personal computers.

[0900] A "server" is a computer system that receives data sent from a terminal and analyzes and processes the data.

[0901] "Means for receiving input" refers to a mechanism by which the terminal can receive text or voice data entered by a user.

[0902] "Means for analyzing input" refers to a mechanism for analyzing received data and understanding its meaning and intent. This includes natural language processing technology.

[0903] "Means for translating" means a mechanism for translating analyzed text into another specified language, including a generative AI model.

[0904] "Means for recognizing emotions" refers to a mechanism for detecting emotions from the user's tone of voice and facial expressions and acquiring that information.

[0905] The "means for adjusting the translation result" refers to a mechanism for correcting the translated text based on the recognized emotional information and adding appropriate emotional expressions.

[0906] "Means for providing the translated results to the user" refers to a mechanism for displaying or audibly notifying the user of the final translation results via a terminal.

[0907] A "generative AI model" is an algorithm that uses artificial intelligence techniques to generate or process data. For example, it includes deep learning models that specialize in natural language processing.

[0908] A "prompt sentence" refers to the text data input into a generative AI model, and is the initial information that the model uses to translate and generate.

[0909] This system translates user-input text and speech into other languages ​​in real time, and also recognizes the user's emotions and reflects them in the translation results. This system is mainly composed of four main components: a server, a terminal, a user, and an emotion recognition engine.

[0910] System configuration and technologies used

[0911] User

[0912] A user accesses their device and inputs text or voice into an input field. For example, a user might input "Hello, what time is it now?" in Japanese. The device then sends the user's tone of voice and facial expressions to an emotion recognition engine.

[0913] Terminal

[0914] The device receives user input in real time and transmits the data and emotion information to the server. The device can be a smartphone or a PC, and data transmission uses the HTTP protocol. The transmitted data includes the input text, user ID, desired language information, and emotion information.

[0915] server

[0916] The server analyzes the data received from the device and translates it into the specified language using a generative AI model. Generative AI models use algorithms specialized for natural language processing, such as OpenAI's GPT-3 model. The server uses an emotion recognition engine (such as Microsoft Azure Cognitive Services) to recognize the user's emotions and adjust the translation results based on that information.

[0917] Emotion Recognition Engine

[0918] The emotion recognition engine analyzes the user's tone of voice and facial expression data to obtain emotional information. Based on this emotional information, the server adjusts the translation results and adds appropriate emotional expressions.

[0919] Specific operation flow

[0920] 1. The user types "Hello, what time is it?". At the same time, the user's excitement level is collected as emotional information.

[0921] 2. The device sends the input and emotion information to the server.

[0922] 3. The server receives the request and parses the data.

[0923] 4. The server uses the generative AI model to translate it as "Hello, what time is it now?"

[0924] 5. The server sends the data to the emotion recognition engine to check the user's excitement level.

[0925] 6. The server adjusts the translation result to "What time is it right now?"

[0926] 7. The server sends the translation results to the device.

[0927] 8. The device displays the translation results to the user.

[0928] Specific examples

[0929] If the user types "Hello, what time is it?" and the tone of voice is detected as excited, the following example prompt will be sent by the device:

[0930] Text input: "Hello, what time is it?"

[0931] Audio Tone Information: "Excited"

[0932] Desired language: "English"

[0933] Based on this, the system generates the translation result "What time is it right now?" and provides it to the user.

[0934] This system enables real-time, emotionally-reflective multilingual communication in the business and tourism sectors, and is equipped with highly accurate error handling capabilities to provide highly reliable services.

[0935] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0936] Program processing steps

[0937] Step 1: User Input

[0938] A user accesses the system and inputs text or voice into the device's input field. The input can include text such as "Hello, what time is it?". Once the input is complete, the device records the voice tone and the user's facial expression and sends it to the emotion recognition engine.

[0939] Input: User-typed text or speech. "Hello, what time is it?"

[0940] Output: Input text, audio, emotional information such as tone of voice and facial expressions.

[0941] Specific action: The user speaks into the smartphone's microphone or enters text on the keyboard.

[0942] Step 2: Submitting input data

[0943] The device sends the text, voice data, and emotion information entered by the user to the server, including the user ID and desired language information.

[0944] Input: User input data and emotional information.

[0945] Output: The data sent to the server as an API request.

[0946] Specific operation: The terminal sends data to the server using the HTTP protocol.

[0947] Step 3: Data analysis

[0948] The server receives the data from the device and first analyzes the user's input, using natural language processing technology to extract the meaning and intent of the text.

[0949] Input: Text data and emotional information sent from the device.

[0950] Output: Parsed semantic and intent information.

[0951] Specific operation: The server retrieves user information from the database and passes the data to the analysis engine.

[0952] Step 4: Translation process

[0953] The server translates the analyzed text into a specified language (e.g., English) using a generative AI model (e.g., GPT-3). The server inputs the prompt sentence into the generative AI model and obtains the translation result.

[0954] Input: Parsed text data.

[0955] Output: Translated text data. "Hello, what time is it now?"

[0956] Specific operation: The server inputs the prompt sentence into the generative AI model and generates the translation result.

[0957] Step 5: Emotion Recognition

[0958] The server uses an emotion recognition engine to recognize the user's emotions based on the transmitted voice tone and facial expression data. The emotion recognition engine returns emotional information such as "excitement" based on the voice tone and facial expression.

[0959] Input: speech tone and facial expression data.

[0960] Output: Recognized emotion information. "Excited"

[0961] Specific operation: The server sends data to the emotion recognition API and obtains emotion information.

[0962] Step 6: Adjust the translation results

[0963] The server adjusts the translation result based on the emotion recognition results. For example, if the user is excited, the server changes the expression to something like "What time is it right now?"

[0964] Input: Translated text data, emotion information.

[0965] Output: Adjusted translation data. "What time is it right now?"

[0966] Specific operation: The server applies logic to reflect emotions in the translation results.

[0967] Step 7: Send the translation

[0968] The server transmits the adjusted translation result to the terminal.

[0969] Input: Adjusted translation data. "What time is it right now?"

[0970] Output: Data sent to the device as an API response.

[0971] Specific operation: The server sends data to the device.

[0972] Step 8: Displaying the translation results

[0973] The device displays the translation results received from the server to the user, and in some cases plays them back aloud.

[0974] Input: Translation data sent from the server.

[0975] Output: The translation result that is displayed to the user.

[0976] Specific operation: The device displays the translation results on the screen and, if necessary, plays the results aloud using a speech synthesis engine.

[0977] (Application example 2)

[0978] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0979] When communicating smoothly with multilingual customers in a brick-and-mortar store or other setting, real-time translation and appropriate reflection of emotions are necessary. However, many conventional systems only translate the language and are unable to provide translation results that reflect the user's emotions. This can lead to communication inaccuracies and emotional gaps, resulting in lower satisfaction and misunderstandings. Technology to resolve this issue is needed.

[0980] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0981] In this invention, the server includes means for receiving input from a user, means for analyzing the received input, means for recognizing the user's emotion, means for translating the analyzed input into a specified language, means for adjusting the translation result based on the emotion, and means for providing the translated result to the user.

[0982] This makes it possible to perform multilingual translation in real time, incorporating the user's emotions, and facilitate smooth communication.

[0983] "User" means any person or entity that accesses and provides input to the System.

[0984] "Input" refers to data that a user provides to a system, either textually or vocally.

[0985] "Analysis" refers to the process of examining and analyzing received input data and preparing it for language translation or emotion recognition.

[0986] The "specified language" refers to the language to which the user wishes to translate, and is a specific language selected from among multiple languages.

[0987] "Translation" refers to the process of converting text or audio expressed in one language into a different language.

[0988] "Emotion" indicates the user's psychological state and includes information recognized from non-verbal elements such as tone of voice and facial expression.

[0989] "Adjustment" refers to changing the translated result to an appropriate expression based on the user's feelings.

[0990] "Terminal" refers to a device used by a user to access the system, such as a smartphone or head-mounted display.

[0991] "Providing" refers to the act of displaying or presenting the processed translation results to the user.

[0992] "Real-time" means that the time between input and translation or adjustment results being provided is very short, with an immediate response.

[0993] This invention is a system that translates text and speech input by a user into other languages ​​in real time, and further recognizes the user's emotions and reflects them in the translation results, facilitating smooth communication. This system is mainly composed of four entities: a server, a terminal, a user, and an emotion engine. Specifically, the system operates as follows:

[0994] Hardware and Software Configuration

[0995] User

[0996] Users access the system through their own devices and input text or voice into the input field. When users input voice, it is converted into text data using voice recognition software (e.g., Google API). At the same time, the user's tone of voice and facial expressions are analyzed to generate emotion data.

[0997] Terminal

[0998] The terminal consists of a device such as a smartphone or a head-mounted display (HMD). The terminal receives input from the user, analyzes the data and emotional information, and sends it to the server. The terminal has a speech recognition library (e.g., the speech_recognition library) and video analysis software installed.

[0999] server

[1000] The server is responsible for analyzing the input data it receives, specifically:

[1001] Translates received text data into a specified language using a generative model (e.g., Hugging Face's transformers library).

[1002] An emotion engine is used to recognize the user's emotions and generate emotion data.

[1003] The translation results are adjusted based on the user's emotional information.

[1004] Emotion Engine

[1005] The emotion engine is a program that recognizes emotions by analyzing the user's tone of voice, facial expressions, etc. The emotion engine is configured using, for example, the sentiment_analysis_toolkit library. This allows it to obtain emotional information such as whether the user is excited or calm.

[1006] Processing Flow

[1007] 1. User Input

[1008] The user speaks, "Hello, what time is it?" At the same time, the user's tone of voice is analyzed to detect excitement.

[1009] 2. Device Operation

[1010] The device converts the voice into text and sends the input data and emotional information to the server.

[1011] 3. Server Processing

[1012] The server analyzes the received input and uses a generative model to translate it as "Hello, what time is it now?". At the same time, it adjusts the translation result to "What time is it right now?" based on emotional information (excitement state).

[1013] 4. Providing translation results

[1014] The server sends the translated and adjusted results to the device, which then displays them to the user.

[1015] This allows users to communicate smoothly in real time with people who speak other languages.

[1016] Specific examples

[1017] For example, if a staff member in a physical store says, "Can I help you with anything?", the voice is recognized and translated into English on the server. Then, based on the user's emotional data, the translation result is adjusted to "Is there anything I can assist you with right now?" and displayed on the staff member's device. This results in smooth communication with the customer.

[1018] Prompt Sentence Examples

[1019] Japanese text: "Hello, what time is it?"

[1020] User Sentiment: "Excited"

[1021] Target language: "English"

[1022] Sample output: "What time is it right now?"

[1023] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1024] Step 1:

[1025] The user speaks, "Hello, what time is it now?" The user's device receives this voice data and converts it into text data using speech recognition software (e.g., Google API). The output is the text data "Hello, what time is it now?"

[1026] Step 2:

[1027] The device analyzes the tone of the user's voice from the audio data and generates emotion data using an emotion engine (e.g., the sentiment_analysis_toolkit library). As a result of the analysis, it is detected that the user is excited, and this is output as emotion information. The output is emotion data for "excited state."

[1028] Step 3:

[1029] The terminal transmits the converted text data and emotion data to the server. The input is the text data "Hello, what time is it now?" and the emotion data "excited."

[1030] Step 4:

[1031] The server parses the received text data and translates it into the specified language (in this case, English) using a generative AI model (for example, Hugging Face's transformers library). The input is the text data "Hello, what time is it now?", and the output is the translation result "Hello, what time is it now?".

[1032] Step 5:

[1033] The server uses an emotion engine to analyze the received emotion data and adjust the translation result. The input is the translation text "Hello, what time is it now?" and the emotion data "excited," and the output is the adjusted translation result of "What time is it right now?"

[1034] Step 6:

[1035] The server sends the adjusted translation result to the terminal. The input is the adjusted translation result of "What time is it right now?" and transfers it to the terminal.

[1036] Step 7:

[1037] The device displays the translation results it receives to the user. The input is the adjusted translation result, "What time is it right now?", which is displayed visually to the user. The output is the message, "What time is it right now?", which the user sees on the device.

[1038] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1039] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1040] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1041] [Fourth embodiment]

[1042] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1043] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1044] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1045] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1046] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1047] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1048] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1049] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1050] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1051] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1052] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1053] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1054] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1055] This invention is a system that translates text and speech input by a user into other languages ​​in real time, facilitating communication. This system is mainly composed of three entities: a server, a terminal, and a user.

[1056] overview

[1057] A user accesses the system and enters text or voice into an input field. For example, a user might enter "Hello, what time is it?"

[1058] The device receives the user's input and sends it to the server, which analyzes the received data and translates it into a specified language (e.g., English) using a generative model.

[1059] Once the translation is complete, the server sends the translation results to the terminal, which then displays them to the user, allowing users to communicate smoothly with users of other languages.

[1060] Specific explanation of each element

[1061] User Input

[1062] Users access the platform through their own devices and enter text or voice into the input fields, which can be used in a variety of situations, such as business meetings or exchanging information at tourist spots.

[1063] For example, consider the case where a user uses their smartphone to type "Hello, what time is it now?" in Japanese.

[1064] Device operation

[1065] The device receives the user's input in real time and transmits the data to the server, for example via its own HTTP protocol, including the user's input text, user ID, and desired language information.

[1066] Server Processing

[1067] The server analyzes the received data. Once the analysis is complete, it uses a generative model to perform the translation process. The generative model utilizes the latest AI technology to achieve high accuracy and real-time translation.

[1068] For example, the server parses the Japanese phrase "Hello, what time is it now?" and translates it into English using the appropriate model, resulting in the translation "Hello, what time is it now?"

[1069] Providing translation results

[1070] After the translation is complete, the server sends the results to the terminal, which then displays them to the user in a user-friendly format, making them instantly available.

[1071] Specific examples

[1072] As a concrete example, the translation flow from Japanese to English is shown below.

[1073] The user types, "Hello, what time is it?"

[1074] The terminal sends the input to the server.

[1075] The server parses the received input and translates it into "Hello, what time is it now?"

[1076] The server sends the translation results to the terminal.

[1077] The terminal displays the translation results to the user.

[1078] This allows users to communicate smoothly with speakers of other languages ​​in real time. This system is particularly expected to be used in the business and tourism sectors, and its error handling functionality will provide a highly reliable service.

[1079] The processing flow will be explained below.

[1080] Step 1:

[1081] A user accesses the platform and enters text or voice into an input field, for example, "Hello, what time is it?"

[1082] Step 2:

[1083] The device receives the user's input and sends the data to the server, including the user's input text, user ID, and desired language information.

[1084] Step 3:

[1085] The server then analyzes the received data, which includes identifying the language of the text entered by the user and performing any necessary pre-processing (e.g., noise removal and text normalization).

[1086] Step 4:

[1087] The server uses the generative model to translate the parsed input data into the specified language. For example, to translate the Japanese phrase "Hello, what time is it now?" into English, it would be converted to "Hello, what time is it now?"

[1088] Step 5:

[1089] The server formats the translation results and sends them to the device, ensuring that the formatted data is organized in a way that is easy for the user to understand.

[1090] Step 6:

[1091] The device receives the translation and displays it to the user, properly formatted in a user-friendly language.

[1092] Step 7:

[1093] The user checks the displayed translation result and takes the next action (re-entering or other operation). If there is further input, the cycle starts again from step 1.

[1094] Through this series of processes, users can communicate smoothly and in real time in other languages.

[1095] Example 1

[1096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1097] Real-time communication between multiple languages ​​is important, especially in business and tourism. However, conventional translation systems often suffer from low translation accuracy and lack real-time capabilities, making them prone to errors. To address this issue, a system that provides high-accuracy, real-time translation and is easy for users to use is needed.

[1098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1099] In this invention, the server includes means for parsing received input, means for using a generative AI model to translate the parsed input into a specified language, and means for providing a prompt sentence to the generative AI model, thereby enabling highly accurate and real-time translation.

[1100] "Input" is text or voice data that a user provides to the system via a terminal.

[1101] "Analysis" is the process by which the server understands the input data it receives and extracts the necessary information.

[1102] A "generative AI model" is an artificial intelligence model that uses technologies such as neural networks to translate input data into a specified language.

[1103] A "prompt sentence" is a sentence that instructs the generative AI model on what processing to do.

[1104] "Translation" is the process of converting input data into a specified language.

[1105] "Providing" is the process in which the server sends the translation result to the terminal, and the terminal displays the result to the user.

[1106] A "terminal" is a device (e.g., a smartphone or tablet) that a user uses to access the system and input data or display results.

[1107] "Server" means a central processing unit that receives input from a terminal and performs analysis, translation, and provides results.

[1108] "Result display" is the process by which the terminal visually presents the translation results received from the server to the user.

[1109] The present invention provides a system for translating text and speech input by a user into other languages ​​in real time to facilitate communication. The system operates in cooperation with a user, a terminal, and a server. A specific embodiment of the system is described below.

[1110] User Input

[1111] Users access the system using devices such as smartphones or tablets. They enter text into the provided input fields or use a microphone to input voice. For example, consider a case where a user enters "Hello, what time is it now?" in Japanese.

[1112] Device operation

[1113] The device receives user input in real time and sends the data to the server. The device formats the user input data and adds necessary metadata such as the user ID and the language to be translated. This data is sent to the server via the HTTP protocol. For example, the device generates an HTTP request containing the text data "Hello, what time is it now?" and sends it to the server.

[1114] Server Processing

[1115] The server receives the request sent from the device and extracts the input text and necessary metadata from the request body. After verifying that the received data is free of errors, it performs the translation using a generative AI model. Specifically, the server generates a prompt sentence for the generative AI model (e.g., OpenAI's GPT-4) and passes it on. An example of a prompt sentence is, "Please translate from Japanese to English. Input sentence: 'Hello, what time is it now?'"

[1116] Providing translation results

[1117] Once the translation process is complete, the server sends the translation result to the terminal as an HTTP response. The response includes the translated text (e.g., "Hello, what time is it now?"). At this time, the response is also checked for formatting and errors.

[1118] Displaying results on your device

[1119] The device analyzes the response received from the server and displays the results to the user. For example, the translation result "Hello, what time is it now?" is displayed on the device's display. The user can check the result in real time.

[1120] Specific examples

[1121] Below is a concrete Japanese to English translation flow:

[1122] 1. The user types "Hello, what time is it?" into the text field.

[1123] 2. The device sends the input to the server along with the "User ID" and "Translation Language: English."

[1124] 3. The server receives the data "Hello, what time is it now?" and checks the content.

[1125] 4. The server sends the generative AI model a prompt: "Please translate from Japanese to English. Input sentence: 'Hello, what time is it now?'" and receives a response from the generative AI model: "Hello, what time is it now?"

[1126] 5. The server sends the translation result, "Hello, what time is it now?" to the terminal.

[1127] 6. The terminal displays the result to the user: "Hello, what time is it now?"

[1128] This system will enable users to communicate smoothly with speakers of other languages ​​in real time, and is expected to be particularly useful in the fields of business and tourism.

[1129] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1130] Step 1: User Input

[1131] A user accesses the system through his / her own terminal. The user inputs text or voice into an input field. Specifically, if the user wants to input "Hello, what time is it now?" in Japanese, he / she types the characters into the input field of the terminal or captures the voice using the microphone. The input data (text or voice) is prepared.

[1132] Step 2: Send data from the device to the server

[1133] The terminal formats the user's input data and adds the necessary metadata (user ID, language to translate, etc.), then sends the formatted request to the server using the HTTP protocol. The input data and metadata are sent to the server.

[1134] Step 3: Data analysis on the server

[1135] The server receives the request from the terminal and analyzes its contents. Specifically, it extracts the input text, user ID, and desired language information from the request body. The server then verifies that the data format and content are correct and performs error checking. The analysis results in structured data.

[1136] Step 4: Translation on the server

[1137] The server performs translation using a generative AI model based on the analysis results. Specifically, it generates a prompt sentence for the generative AI model (e.g., generative model GPT-4) and passes it on. For example, a prompt sentence such as "Please translate from Japanese to English. Input sentence: 'Hello, what time is it now?'" is generated and input into the AI ​​model. The generative AI model converts the input data into the specified language and outputs the translated data, "Hello, what time is it now?"

[1138] Step 5: Send the translation results from the server to the device

[1139] The server formats the translation results obtained from the generative AI model as an HTTP response and sends it to the device. The response includes the translation result (e.g., "Hello, what time is it now?"). The server checks the response format and content for errors. The formatted response data is sent to the device.

[1140] Step 6: View the results in your device

[1141] The device analyzes the response received from the server and extracts the translation result. Specifically, the device extracts the translated text from the response data and displays it to the user. The displayed result (e.g., "Hello, what time is it now?") is confirmed by the user.

[1142] Through these steps, users can obtain highly accurate translation results into other languages ​​in real time.

[1143] (Application example 1)

[1144] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1145] In modern society, situations requiring multilingual communication in brick-and-mortar stores are on the rise, but language barriers can make smooth customer service difficult. This problem is particularly pronounced in regions where the number of foreign language speakers is increasing, requiring accurate translation in real time. However, existing translation systems often lack sufficient translation speed and accuracy. Furthermore, they lack the ability to check history, making it impossible to refer to past communications, which risks reducing the quality of service.

[1146] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1147] In this invention, the server includes means for receiving input from a user, means for analyzing the received input, means for translating the analyzed input into a specified language, means for providing the translated result to the user, and means for providing an interface for users and foreign language speakers to communicate in real time at the store. This enables accurate multilingual communication in real time. The inclusion of a history function also allows past communications to be referenced, improving the quality of service. Furthermore, translation is performed using a generative model, and accuracy is improved by specific prompt sentences, providing more reliable translation results.

[1148] "User" refers to a person who uses the system to input text or voice and receive the results translated into another language.

[1149] "Input" refers to the text or voice information provided by a user to a system.

[1150] "Analysis" refers to the information processing means that processes received input data and makes sense of it.

[1151] "Translation" refers to the process of converting parsed input data into another specified language.

[1152] "Interface" refers to the operation screen and operation means that facilitate smooth interaction between the user and the system.

[1153] The "history function" refers to a function that saves past communication details with users and allows them to refer to them as needed.

[1154] "Generative model" refers to a machine learning algorithm that uses AI technology to produce highly accurate translations.

[1155] A "prompt sentence" refers to an instruction sentence given to a generative model to perform an appropriate translation.

[1156] To implement this invention, it is necessary to build a system in which users, terminals, and servers work together to provide highly accurate real-time translation and support smooth communication between multiple languages.

[1157] User operations

[1158] Users access the system using devices such as smartphones or tablets. The device displays an interface as an operation screen, where users input text or voice. For example, consider the case where a store clerk asks a foreign customer, "Hello, how do you want to use this product?" in Japanese.

[1159] Device operation

[1160] The terminal receives text and voice input from the user and sends it to the server, generating a data packet containing the user's input data, user ID, and desired language information using a communication method such as the HTTP protocol.

[1161] Server Processing

[1162] The server analyzes the input data sent from the device. After analysis, it uses a generative AI model to translate it into the specified language. Specifically, it uses OpenAI's API to translate based on the following prompt:

[1163] Prompt Sentence Examples

[1164] "Translate this text to English: Hello, how do I use this product?"

[1165] Based on this prompt, the generative model generates a translation result such as "Hello, how do I use this product?"

[1166] Providing translation results

[1167] The server then sends the translation result back to the terminal, which then displays it in a format that is easy for the user to understand. For example, the clerk's terminal screen might display "Hello, how do I use this product?" in English.

[1168] History function

[1169] Furthermore, the system has a built-in history function that allows past communications to be saved, allowing users to refer to past conversation history as needed to improve the quality of service.

[1170] Hardware and Software

[1171] Hardware: Smartphone or tablet (iOS or Android)

[1172] Software: The front end uses HTML, CSS, and JavaScript, and the back end uses Python Flask and the OpenAI API.

[1173] By combining these steps, the system facilitates multilingual communication in brick-and-mortar stores, enabling users and foreign language speakers to communicate effectively in real time.

[1174] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1175] Step 1:

[1176] A user accesses the system using a device such as a smartphone or tablet and inputs text or voice into an input field. For example, a store clerk might input "Hello, how do you want to use this product?" in Japanese. In this case, the input is sent to the device as text or voice data.

[1177] Step 2:

[1178] The device analyzes the received text and voice data and sends it to the server. At this point, a data packet is formed using the HTTP protocol or similar, and information such as the user ID and desired language (e.g., English) is also included. This is the process by which the text and voice data received as input is converted into a data packet to be sent to the server.

[1179] Step 3:

[1180] The server receives data packets sent from the device and analyzes the input data. After analysis, it uses a generative AI model to translate it into the specified language. For example, the server provides the prompt sentence "Translate this text to English: Hello, how do you use this product?" to the generative model, and the resulting translation is "Hello, how do I use this product?"

[1181] Step 4:

[1182] The server then sends the translation results back to the terminal. This time, the translation results are sent in text format. The server's input is the translated text data, and it is the process that processes it to send it to the terminal.

[1183] Step 5:

[1184] The terminal receives the translation results from the server and displays them to the user. The results are displayed visually through a user-friendly interface, allowing the user to understand them immediately. For example, the salesperson's terminal screen might display "Hello, how do I use this product?"

[1185] Step 6:

[1186] As a history function, the device will store past translation history and allow users to refer to this history as needed, supporting continuous communication with the same user and improving the quality of service. The history data will be stored in local storage or cloud storage, allowing users to access it at any time.

[1187] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1188] This invention is a system that translates user-entered text and speech into other languages ​​in real time, and also recognizes the user's emotions and reflects them in the translation results, facilitating smooth communication. This system is mainly composed of four main components: a server, a terminal, a user, and an emotion engine.

[1189] overview

[1190] A user accesses the system and inputs text or voice into an input field. For example, consider the case where a user inputs "Hello, what time is it now?". At the same time, the system reads the user's emotions from their tone of voice and facial expressions.

[1191] The device receives user input and sends the data and emotion information to the server, which analyzes the received data and translates it into a specified language (e.g., English) using a generative model. It also uses an emotion engine to recognize the user's emotion and adjust the translation results based on that emotion.

[1192] Once the translation is complete, the server sends the translation results to the terminal, which then displays them to the user, allowing users to communicate smoothly with users of other languages.

[1193] Specific explanation of each element

[1194] User Input

[1195] Users access the platform through their own devices and enter text or voice into the input field. This is used in a variety of situations, such as business meetings and exchanging information at tourist spots. For example, consider a case where a user uses their smartphone to enter "Hello, what time is it now?" in Japanese.

[1196] Device operation

[1197] The device receives user input in real time and transmits the data and emotion information to the server via its own HTTP protocol, etc. The transmitted data includes the user's input text, user ID, desired language information, and emotion information.

[1198] Server Processing

[1199] The server analyzes the received data. Once the analysis is complete, it uses a generative model to perform the translation process. The generative model utilizes the latest AI technology to achieve high accuracy and real-time translation.

[1200] For example, the server analyzes the Japanese phrase "Hello, what time is it now?" and translates it into English using an appropriate model, resulting in the translation "Hello, what time is it now?"

[1201] In parallel, the server uses an emotion engine to recognize the user's emotions and adjusts the translated text accordingly. For example, if the user is excited, the server might translate it as "What time is it right now?"

[1202] Providing translation results

[1203] After the translation is complete, the server sends the results to the device, which then displays them to the user in a user-friendly format, taking into account emotional information.

[1204] Specific examples

[1205] As a concrete example, the translation flow from Japanese to English is shown below.

[1206] A user types, "Hello, what time is it?" At the same time, the user's tone of voice is analyzed and it is detected that the user is excited.

[1207] The device sends the input and emotion information to the server.

[1208] The server analyzes the received input and translates it as "Hello, what time is it now?". It also adjusts the translation result based on the emotion information to "What time is it right now?"

[1209] The server sends the translation results to the terminal.

[1210] The terminal displays the translation results to the user.

[1211] This allows users to communicate smoothly and emotionally with speakers of other languages ​​in real time. This system is particularly expected to be used in the business and tourism sectors, and its error handling functionality will provide a highly reliable service.

[1212] The processing flow will be explained below.

[1213] Step 1:

[1214] A user accesses the platform and enters text or voice into an input field, for example, "Hello, what time is it?" The user's facial expressions and tone of voice are also recorded.

[1215] Step 2:

[1216] The device receives user input and converts it into text data, while also collecting the user's emotional information (e.g., excitement, sadness, joy).

[1217] Step 3:

[1218] The device sends the user's input data and emotion information to the server, along with the user ID and desired language.

[1219] Step 4:

[1220] The server analyzes the received data, which includes identifying the language of the text and pre-processing it (denoising, normalization).

[1221] Step 5:

[1222] The server uses the generative model to translate the parsed text into the specified language. For example, the Japanese phrase "Hello, what time is it now?" is translated into English as "Hello, what time is it now?"

[1223] Step 6:

[1224] The server uses the emotion engine to recognize the user's emotion. For example, if the user is excited, the emotion is recognized as "excited."

[1225] Step 7:

[1226] The server adjusts the translation results based on the emotional information it recognizes. For example, if the user is feeling excited, it will modify "Hello, what time is it now?" to reflect that emotion, such as "What time is it right now?"

[1227] Step 8:

[1228] The server formats the final translation and sends it to the device, with adjustments based on emotional information.

[1229] Step 9:

[1230] The device will then display the translation results to the user, specifically asking, "What time is it right now?"

[1231] Step 10:

[1232] The user checks the displayed translation result and takes the next action (re-entering or other operations) if necessary. If re-entering is required, the cycle starts again from step 1.

[1233] Through this series of processes, users can communicate smoothly and emotionally with speakers of other languages ​​in real time.

[1234] Example 2

[1235] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1236] In modern society, real-time communication between people who speak different languages ​​is becoming increasingly important. However, conventional translation systems simply translate text and are unable to provide translation results that reflect the user's emotions. This makes it difficult to accurately convey important emotional information in communication.

[1237] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving input from a user, means for analyzing the received input, means for translating the analyzed input into a specified language, means for recognizing the user's emotion, means for adjusting the translation result based on the emotion recognition result, and means for providing the translation result to the user. This makes it possible to provide a translation result that reflects emotional information contained in the text or voice input by the user.

[1238] "User" means any individual or entity that uses the System.

[1239] "Terminal" refers to an electronic device that allows a user to access and operate the system. Examples include smartphones and personal computers.

[1240] A "server" is a computer system that receives data sent from a terminal and analyzes and processes the data.

[1241] "Means for receiving input" refers to a mechanism by which the terminal can receive text or voice data entered by a user.

[1242] "Means for analyzing input" refers to a mechanism for analyzing received data and understanding its meaning and intent. This includes natural language processing technology.

[1243] "Means for translating" means a mechanism for translating analyzed text into another specified language, including a generative AI model.

[1244] "Means for recognizing emotions" refers to a mechanism for detecting emotions from the user's tone of voice and facial expressions and acquiring that information.

[1245] The "means for adjusting the translation result" refers to a mechanism for correcting the translated text based on the recognized emotional information and adding appropriate emotional expressions.

[1246] "Means for providing the translated results to the user" refers to a mechanism for displaying or audibly notifying the user of the final translation results via a terminal.

[1247] A "generative AI model" is an algorithm that uses artificial intelligence techniques to generate or process data. For example, it includes deep learning models that specialize in natural language processing.

[1248] A "prompt sentence" refers to the text data input into a generative AI model, and is the initial information that the model uses to translate and generate.

[1249] This system translates user-input text and speech into other languages ​​in real time, and also recognizes the user's emotions and reflects them in the translation results. This system is mainly composed of four main components: a server, a terminal, a user, and an emotion recognition engine.

[1250] System configuration and technologies used

[1251] User

[1252] A user accesses their device and inputs text or voice into an input field. For example, a user might input "Hello, what time is it now?" in Japanese. The device then sends the user's tone of voice and facial expressions to an emotion recognition engine.

[1253] Terminal

[1254] The device receives user input in real time and transmits the data and emotion information to the server. The device can be a smartphone or a PC, and data transmission uses the HTTP protocol. The transmitted data includes the input text, user ID, desired language information, and emotion information.

[1255] server

[1256] The server analyzes the data received from the device and translates it into the specified language using a generative AI model. Generative AI models use algorithms specialized for natural language processing, such as OpenAI's GPT-3 model. The server uses an emotion recognition engine (such as Microsoft Azure Cognitive Services) to recognize the user's emotions and adjust the translation results based on that information.

[1257] Emotion Recognition Engine

[1258] The emotion recognition engine analyzes the user's tone of voice and facial expression data to obtain emotional information. Based on this emotional information, the server adjusts the translation results and adds appropriate emotional expressions.

[1259] Specific operation flow

[1260] 1. The user types "Hello, what time is it?". At the same time, the user's excitement level is collected as emotional information.

[1261] 2. The device sends the input and emotion information to the server.

[1262] 3. The server receives the request and parses the data.

[1263] 4. The server uses the generative AI model to translate it as "Hello, what time is it now?"

[1264] 5. The server sends the data to the emotion recognition engine to check the user's excitement level.

[1265] 6. The server adjusts the translation result to "What time is it right now?"

[1266] 7. The server sends the translation results to the device.

[1267] 8. The device displays the translation results to the user.

[1268] Specific examples

[1269] If the user types "Hello, what time is it?" and the tone of voice is detected as excited, the following example prompt will be sent by the device:

[1270] Text input: "Hello, what time is it?"

[1271] Audio Tone Information: "Excited"

[1272] Desired language: "English"

[1273] Based on this, the system generates the translation result "What time is it right now?" and provides it to the user.

[1274] This system enables real-time, emotionally-reflective multilingual communication in the business and tourism sectors, and is equipped with highly accurate error handling capabilities to provide highly reliable services.

[1275] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1276] Program processing steps

[1277] Step 1: User Input

[1278] A user accesses the system and inputs text or voice into the device's input field. The input can include text such as "Hello, what time is it?". Once the input is complete, the device records the voice tone and the user's facial expression and sends it to the emotion recognition engine.

[1279] Input: User-typed text or speech. "Hello, what time is it?"

[1280] Output: Input text, audio, emotional information such as tone of voice and facial expressions.

[1281] Specific action: The user speaks into the smartphone's microphone or enters text on the keyboard.

[1282] Step 2: Submitting input data

[1283] The device sends the text, voice data, and emotion information entered by the user to the server, including the user ID and desired language information.

[1284] Input: User input data and emotional information.

[1285] Output: The data sent to the server as an API request.

[1286] Specific operation: The terminal sends data to the server using the HTTP protocol.

[1287] Step 3: Data analysis

[1288] The server receives the data from the device and first analyzes the user's input, using natural language processing technology to extract the meaning and intent of the text.

[1289] Input: Text data and emotional information sent from the device.

[1290] Output: Parsed semantic and intent information.

[1291] Specific operation: The server retrieves user information from the database and passes the data to the analysis engine.

[1292] Step 4: Translation process

[1293] The server translates the analyzed text into a specified language (e.g., English) using a generative AI model (e.g., GPT-3). The server inputs the prompt sentence into the generative AI model and obtains the translation result.

[1294] Input: Parsed text data.

[1295] Output: Translated text data. "Hello, what time is it now?"

[1296] Specific operation: The server inputs the prompt sentence into the generative AI model and generates the translation result.

[1297] Step 5: Emotion Recognition

[1298] The server uses an emotion recognition engine to recognize the user's emotions based on the transmitted voice tone and facial expression data. The emotion recognition engine returns emotional information such as "excitement" based on the voice tone and facial expression.

[1299] Input: speech tone and facial expression data.

[1300] Output: Recognized emotion information. "Excited"

[1301] Specific operation: The server sends data to the emotion recognition API and obtains emotion information.

[1302] Step 6: Adjust the translation results

[1303] The server adjusts the translation result based on the emotion recognition results. For example, if the user is excited, the server changes the expression to something like "What time is it right now?"

[1304] Input: Translated text data, emotion information.

[1305] Output: Adjusted translation data. "What time is it right now?"

[1306] Specific operation: The server applies logic to reflect emotions in the translation results.

[1307] Step 7: Send the translation

[1308] The server transmits the adjusted translation result to the terminal.

[1309] Input: Adjusted translation data. "What time is it right now?"

[1310] Output: Data sent to the device as an API response.

[1311] Specific operation: The server sends data to the device.

[1312] Step 8: Displaying the translation results

[1313] The device displays the translation results received from the server to the user, and in some cases plays them back aloud.

[1314] Input: Translation data sent from the server.

[1315] Output: The translation result that is displayed to the user.

[1316] Specific operation: The device displays the translation results on the screen and, if necessary, plays the results aloud using a speech synthesis engine.

[1317] (Application example 2)

[1318] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1319] When communicating smoothly with multilingual customers in a brick-and-mortar store or other setting, real-time translation and appropriate reflection of emotions are necessary. However, many conventional systems only translate the language and are unable to provide translation results that reflect the user's emotions. This can lead to communication inaccuracies and emotional gaps, resulting in lower satisfaction and misunderstandings. Technology to resolve this issue is needed.

[1320] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1321] In this invention, the server includes means for receiving input from a user, means for analyzing the received input, means for recognizing the user's emotion, means for translating the analyzed input into a specified language, means for adjusting the translation result based on the emotion, and means for providing the translated result to the user.

[1322] This makes it possible to perform multilingual translation in real time, incorporating the user's emotions, and facilitate smooth communication.

[1323] "User" means any person or entity that accesses and provides input to the System.

[1324] "Input" refers to data that a user provides to a system, either textually or vocally.

[1325] "Analysis" refers to the process of examining and analyzing received input data and preparing it for language translation or emotion recognition.

[1326] The "specified language" refers to the language to which the user wishes to translate, and is a specific language selected from among multiple languages.

[1327] "Translation" refers to the process of converting text or audio expressed in one language into a different language.

[1328] "Emotion" indicates the user's psychological state and includes information recognized from non-verbal elements such as tone of voice and facial expression.

[1329] "Adjustment" refers to changing the translated result to an appropriate expression based on the user's feelings.

[1330] "Terminal" refers to a device used by a user to access the system, such as a smartphone or head-mounted display.

[1331] "Providing" refers to the act of displaying or presenting the processed translation results to the user.

[1332] "Real-time" means that the time between input and translation or adjustment results being provided is very short, with an immediate response.

[1333] This invention is a system that translates text and speech input by a user into other languages ​​in real time, and further recognizes the user's emotions and reflects them in the translation results, facilitating smooth communication. This system is mainly composed of four entities: a server, a terminal, a user, and an emotion engine. Specifically, the system operates as follows:

[1334] Hardware and Software Configuration

[1335] User

[1336] Users access the system through their own devices and input text or voice into the input field. When users input voice, it is converted into text data using voice recognition software (e.g., Google API). At the same time, the user's tone of voice and facial expressions are analyzed to generate emotion data.

[1337] Terminal

[1338] The terminal consists of a device such as a smartphone or a head-mounted display (HMD). The terminal receives input from the user, analyzes the data and emotional information, and sends it to the server. The terminal has a speech recognition library (e.g., the speech_recognition library) and video analysis software installed.

[1339] server

[1340] The server is responsible for analyzing the input data it receives, specifically:

[1341] Translates received text data into a specified language using a generative model (e.g., Hugging Face's transformers library).

[1342] An emotion engine is used to recognize the user's emotions and generate emotion data.

[1343] The translation results are adjusted based on the user's emotional information.

[1344] Emotion Engine

[1345] The emotion engine is a program that recognizes emotions by analyzing the user's tone of voice, facial expressions, etc. The emotion engine is configured using, for example, the sentiment_analysis_toolkit library. This allows it to obtain emotional information such as whether the user is excited or calm.

[1346] Processing Flow

[1347] 1. User Input

[1348] The user speaks, "Hello, what time is it?" At the same time, the user's tone of voice is analyzed to detect excitement.

[1349] 2. Device Operation

[1350] The device converts the voice into text and sends the input data and emotional information to the server.

[1351] 3. Server Processing

[1352] The server analyzes the received input and uses a generative model to translate it as "Hello, what time is it now?". At the same time, it adjusts the translation result to "What time is it right now?" based on emotional information (excitement state).

[1353] 4. Providing translation results

[1354] The server sends the translated and adjusted results to the device, which then displays them to the user.

[1355] This allows users to communicate smoothly in real time with people who speak other languages.

[1356] Specific examples

[1357] For example, if a staff member in a physical store says, "Can I help you with anything?", the voice is recognized and translated into English on the server. Then, based on the user's emotional data, the translation result is adjusted to "Is there anything I can assist you with right now?" and displayed on the staff member's device. This results in smooth communication with the customer.

[1358] Prompt Sentence Examples

[1359] Japanese text: "Hello, what time is it?"

[1360] User Sentiment: "Excited"

[1361] Target language: "English"

[1362] Sample output: "What time is it right now?"

[1363] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1364] Step 1:

[1365] The user speaks, "Hello, what time is it now?" The user's device receives this voice data and converts it into text data using speech recognition software (e.g., Google API). The output is the text data "Hello, what time is it now?"

[1366] Step 2:

[1367] The device analyzes the tone of the user's voice from the audio data and generates emotion data using an emotion engine (e.g., the sentiment_analysis_toolkit library). As a result of the analysis, it is detected that the user is excited, and this is output as emotion information. The output is emotion data for "excited state."

[1368] Step 3:

[1369] The terminal transmits the converted text data and emotion data to the server. The input is the text data "Hello, what time is it now?" and the emotion data "excited."

[1370] Step 4:

[1371] The server parses the received text data and translates it into the specified language (in this case, English) using a generative AI model (for example, Hugging Face's transformers library). The input is the text data "Hello, what time is it now?", and the output is the translation result "Hello, what time is it now?".

[1372] Step 5:

[1373] The server uses an emotion engine to analyze the received emotion data and adjust the translation result. The input is the translation text "Hello, what time is it now?" and the emotion data "excited," and the output is the adjusted translation result of "What time is it right now?"

[1374] Step 6:

[1375] The server sends the adjusted translation result to the terminal. The input is the adjusted translation result of "What time is it right now?" and transfers it to the terminal.

[1376] Step 7:

[1377] The device displays the translation results it receives to the user. The input is the adjusted translation result, "What time is it right now?", which is displayed visually to the user. The output is the message, "What time is it right now?", which the user sees on the device.

[1378] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1379] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1380] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1381] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1382] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1383] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1384] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1385] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, motorcycles, and other devices, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1386] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1387] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1388] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1389] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1390] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1391] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1392] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1393] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1394] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1395] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1396] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1397] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1398] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1399] The following is further disclosed regarding the above embodiment.

[1400] (Claim 1)

[1401] means for receiving input from a user;

[1402] means for analyzing the received input;

[1403] means for translating the parsed input into a specified language;

[1404] means for providing the translated result to a user.

[1405] (Claim 2)

[1406] 10. The system of claim 1, wherein the user input is text or voice.

[1407] (Claim 3)

[1408] 2. The system according to claim 1, wherein the translation means performs the translation using a generative model.

[1409] (Claim 4)

[1410] 2. The system according to claim 1, wherein the providing means provides the translation result to the user in real time.

[1411] (Claim 5)

[1412] 2. The system of claim 1, wherein the analyzing means analyzes the input data along with the user's ID and intended language information.

[1413] (Claim 6)

[1414] The system of claim 1, wherein the system is used in the field of business or tourism.

[1415] (Claim 7)

[1416] 2. The system of claim 1, wherein said system comprises an error handling means for performing error checking at each step of processing.

[1417] "Example 1"

[1418] (Claim 1)

[1419] means for receiving input from a user;

[1420] means for analyzing the received input;

[1421] means for using a generative AI model to translate the parsed input into a specified language;

[1422] means for providing a prompt sentence to the generative AI model;

[1423] means for providing the translated result to a user;

[1424] said user input being text or voice;

[1425] The system further includes means for displaying the translation result on a user's terminal.

[1426] (Claim 2)

[1427] The system according to claim 1, wherein the server performs the translation process using the generative AI model and transmits the translation result to a terminal.

[1428] (Claim 3)

[1429] 2. The system according to claim 1, wherein the terminal analyzes the translation result received from the server and displays it to the user.

[1430] "Application Example 1"

[1431] (Claim 1)

[1432] means for receiving input from a user;

[1433] means for analyzing the received input;

[1434] means for translating the parsed input into a specified language;

[1435] means for providing the translated result to a user; and

[1436] A system including means for providing an interface for users and foreign language speakers to communicate in real time in a store.

[1437] (Claim 2)

[1438] 10. The system of claim 1, wherein the user's input is text or voice, and further comprising means for displaying the translation results and a history function for facilitating the real-time communication.

[1439] (Claim 3)

[1440] The system of claim 1 , wherein the translation means uses a generative model to perform the translation and refines it with specific prompt sentences.

[1441] "Example 2: Combining Emotion Engines"

[1442] (Claim 1)

[1443] means for receiving input from a user;

[1444] means for analyzing the received input;

[1445] means for translating the parsed input into a specified language;

[1446] means for recognizing a user's emotion;

[1447] means for adjusting the translation result based on the emotion recognition result;

[1448] means for providing the translated result to a user.

[1449] (Claim 2)

[1450] 10. The system of claim 1, wherein the user input is text or voice.

[1451] (Claim 3)

[1452] 2. The system of claim 1, wherein the translation means performs the translation using a generative AI model.

[1453] "Application example 2 when combining emotion engines"

[1454] (Claim 1)

[1455] means for receiving input from a user;

[1456] means for analyzing the received input;

[1457] means for translating the parsed input into a specified language;

[1458] means for recognizing a user's emotion;

[1459] means for adjusting a translation result based on said emotion;

[1460] means for providing the translated result to a user.

[1461] (Claim 2)

[1462] 10. The system of claim 1, wherein the user input is text or voice.

[1463] (Claim 3)

[1464] 2. The system according to claim 1, wherein the translation means performs the translation using a generative model.

[1465] (Claim 4)

[1466] 10. The system of claim 1, wherein the received input is received via a terminal.

[1467] (Claim 5)

[1468] 10. The system according to claim 1, further comprising means for analyzing the user's emotions using tone of voice and facial expressions.

[1469] (Claim 6)

[1470] The system of claim 1, further comprising means for displaying the adjusted translation result on a smartphone or a head-mounted display. [Explanation of symbols]

[1471] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving input from a user; means for analyzing the received input; means for translating the parsed input into a specified language; means for providing the translated result to a user.

2. The system of claim 1 , wherein the user input is text or voice.

3. The system according to claim 1 , wherein the translation means performs the translation using a generative model.

4. 2. The system according to claim 1, wherein said providing means provides the translation result to the user in real time.

5. 2. The system of claim 1, wherein said analyzing means analyzes the input data along with user ID and intended language information.

6. The system of claim 1, wherein the system is used in the field of business or tourism.

7. 2. The system of claim 1, wherein said system includes an error handling means for performing error checking at each step of processing.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A