system

A system automatically detects and translates display languages using generative AI, addressing the lack of multilingual support in websites and applications, enhancing user experience for foreign language speakers.

JP2026104477APending Publication Date: 2026-06-25SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-13
Publication Date
2026-06-25

Smart Images

  • Figure 2026104477000001_ABST
    Figure 2026104477000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】 A device for the terminal to save the language setting selected by the user, A device for the terminal to detect the display language of the information source being accessed, A device for the terminal to generate a translation request based on the detected language and the language selected by the user and transmit it to the communication device, A device for the communication device to translate data based on the received translation request, A device for the communication device to transmit the translation result to the terminal, A device for the terminal to display the received translation result, A device for the terminal to obtain the information presented using the identification image and generate a translation request, A system including
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] There are many websites and applications with insufficient multilingual support, making it difficult for foreign language speakers to obtain information quickly and appropriately. Also, on the side implementing multilingual support, there is a problem of a large burden on manpower and time. Especially in tourist destinations and public transportation, multilingual guidance is required, but the response is delayed, resulting in the current situation where foreign visitors to Japan feel inconvenienced.

Means for Solving the Problems

[0005] This invention provides a means for a terminal to save the user's language settings and automatically detect the display language of websites and applications. Furthermore, it enables multilingual support by having the terminal send translation requests it generates to a server, which then translates the text using AI. The translation results are returned to the terminal and displayed in a specific part of the user interface, allowing the user to obtain information in their own language. This reduces the time and effort required for multilingual support and significantly improves convenience for foreign language speakers.

[0006] A "device" refers to a device used by a user, such as a smartphone, tablet, or computer.

[0007] "User-selected language" refers to the language setting for information specified by the user on the device.

[0008] "Display language" refers to the language in which the text used within a website or application is written.

[0009] A "translation request" is a request sent from a terminal to a server to translate specific text into a specified language.

[0010] A "server" refers to a computer system that can receive, process, and send back data via a network.

[0011] "Generative AI" refers to a technology that uses artificial intelligence, specifically a language model trained to translate specified text.

[0012] "Translation result" refers to the translated text data that the server returns to the terminal in the specified language.

[0013] "User interface" refers to the screens and operational functions that allow the user and the application to interact on a device.

[0014] "Multilingual support" means that a single system or service supports multiple languages ​​and provides information accordingly. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a processor with a reference number (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a RAM (Random Access Memory) with a reference number is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a storage with a reference number is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is a multilingual system utilizing a terminal, a server, and a generative AI. The terminal used by the user simultaneously launches its own translation agent when an application or website is started. This translation agent records the language selected by the user on the terminal as initial setting data and automatically detects the display language of the content the terminal is accessing.

[0037] Based on the detected display language, the terminal generates and sends a translation request to the server. This translation request includes the original language, the target language (the language selected by the user), and the text information that needs to be translated.

[0038] The server processes the received request and forwards it to the generative AI. The generative AI uses advanced language models, such as neural network models, to translate the received text into the user's selected language. The translated text information is then sent to the terminal by the server as the translation result.

[0039] The device translates and displays the relevant parts of the application or website's user interface based on the translation results received from the server. This allows the user to view the information in their chosen language.

[0040] As a concrete example, consider a scenario where a user who only speaks English uses a Japanese public transportation timetable app. The user launches the app and sets English as the preferred language. The device's translation agent detects the timetable information provided in Japanese and sends a request to the server to translate it into English. The server uses generative AI to translate the information, and the translated timetable is displayed on the device. Based on this, the user can plan their travel schedule and enjoy a comfortable sightseeing experience.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The user launches the application on their device and sets their preferred language. This setting information is saved on the device and used in subsequent processes.

[0044] Step 2:

[0045] The device activates a translation agent and detects the display language of the content the user is viewing. This detection is performed by analyzing the text on the page.

[0046] Step 3:

[0047] The device creates a translation request based on the detected display language and the language selected by the user, and sends it to the server. This request also includes the specific text that needs to be translated.

[0048] Step 4:

[0049] The server processes the received translation request and passes it to the generating AI. The AI ​​model translates the requested text into the user's selected language.

[0050] Step 5:

[0051] The server sends the translated result to the terminal. This result is data containing the translated text.

[0052] Step 6:

[0053] The device translates the user interface of an application or website based on the received translation results and displays it to the user.

[0054] Step 7:

[0055] Users can perform the necessary actions using translated content, thereby enabling them to accurately understand and utilize the information.

[0056] (Example 1)

[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0058] When accessing information in different languages, users often face challenges in understanding the information, particularly in the interface, which can make smooth operation difficult. This means that language barriers are a significant obstacle to information utilization and international communication.

[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0060] In this invention, the server includes means for converting character data based on a received translation request, means for transmitting the conversion result to a terminal, and means for receiving translation requests from the terminal. This enables smooth translation and display of information in different languages.

[0061] A "terminal" is an electronic device that users directly operate to access information.

[0062] "Language settings" is a component that allows users to specify their preferred language for displaying and interacting with information.

[0063] "Display language" refers to the language used when information or an interface is presented visually to the user.

[0064] "Means of recognition" refers to a function that allows the terminal to automatically identify the language in which information is displayed.

[0065] A "translation request" is a request sent from a terminal to a server that includes the source language and the target language of the information.

[0066] A "device for aggregating information" is a central computer that performs data conversion and management based on requests from terminals.

[0067] "Means of converting text data" refers to the process of automatically translating information into another language.

[0068] "Conversion result" refers to the output after converting character data to another language.

[0069] "User interface" refers to the screen that the user directly interacts with, where information is presented through the terminal.

[0070] "Generative artificial intelligence" is a technology that uses neural network models to automatically translate text into other languages.

[0071] This invention is a system that facilitates multilingual information access and includes a terminal for user access, a server for aggregating information, and a generative AI with advanced translation capabilities.

[0072] The user launches a specific source of information (e.g., a webpage or application) on their device. This device has a feature that allows the user to specify their preferred display language as the default setting. The device automatically parses HTML tags and metadata to recognize the language of the displayed information.

[0073] The device generates a translation request to a server that aggregates information, based on its recognized language settings. This request includes the original language, the user's desired target language, and the text to be translated. The server receives this request and sends it as a prompt to the artificial intelligence that generates the information.

[0074] The generative AI uses a neural network model to translate text into the desired language. This AI considers context and cultural nuances to produce natural-sounding translations. The translated text is then sent to the terminal via a server.

[0075] The terminal displays the received conversion result in the user's preferred language on the interface they interact with. This allows the user to understand and interact with information in their chosen language.

[0076] As a concrete example, consider a scenario where a user who only understands English accesses a tourist information webpage provided in Japanese. In this case, the user opens the webpage and specifies English as the desired display language. The device recognizes the Japanese content and sends a request to the server to translate it into English. The server uses generative AI to translate the information, and the English content is displayed on the device. This allows the user to use the information without being aware of the language barrier.

[0077] An example of a prompt message would be the text format: "Please translate the following Japanese sentence into English: 'Hello'".

[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0079] Step 1:

[0080] The user launches a specific application or website using the device. The translation agent on the device records the language the user initially selected. Based on this input, the device retains the user's preferred language as output.

[0081] Step 2:

[0082] The device automatically detects the display language of the content being accessed. This process involves analyzing HTML tags and metadata, for example, by checking the "lang" attribute to identify the language. The detected language is then used as input data for the translation process.

[0083] Step 3:

[0084] The device generates a translation request based on the detected display language and the user's language setting, and sends it to the server. This translation request includes the source language, the target language, and the text to be translated. This is then sent to the server as output data for the translation process.

[0085] Step 4:

[0086] The server analyzes the translation request received from the terminal and forwards the request to the generating AI model along with the generated prompt. The server's input is the translation request, and its output is the AI ​​request, which includes the prompt.

[0087] Step 5:

[0088] The generative AI processes data based on the received prompt to produce a natural translation. The AI ​​uses a neural network to analyze context and nuances, and outputs the translated text. The output is the translated text in the language selected by the user.

[0089] Step 6:

[0090] The server receives the translated text from the generating AI and sends it to the user's terminal. The server's input is the AI's translation result, and its output is the converted data sent to the terminal.

[0091] Step 7:

[0092] The terminal translates the necessary parts of the user interface based on the received translated text. This allows the user to view the information in their preferred language. The input is the translation result from the server, and the output is the translated user interface.

[0093] (Application Example 1)

[0094] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0095] In a multilingual environment, a key challenge is to mitigate the difficulties users face due to language barriers when purchasing or paying for goods, enabling them to conduct business transactions with confidence. In particular, when visiting a country or region where different languages ​​are spoken, it is essential to quickly and accurately translate information into the user's native language and provide smooth service.

[0096] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0097] In this invention, the server includes a device that stores the user's selected language settings on the terminal, a device that acquires information presented using an identification image on the terminal and generates a translation request, and a device that performs the translation process using the generated artificial intelligence. This enables users to quickly verify information in their native language and conduct transactions reliably, even in different language environments.

[0098] A "terminal" is an electronic device that a user directly operates to display information.

[0099] A "user" is defined as the entity that uses the system to translate or verify information.

[0100] "Language settings" means that the user specifies their native language or preferred language.

[0101] A "device" is a component of hardware or software designed to perform a specific function.

[0102] "Information sources" refer to websites and applications accessible through a device.

[0103] "Detection" is the act of analyzing information in order to identify a specified language.

[0104] A "translation request" is the process of requesting that information be converted from the original language to the user's desired language.

[0105] A "communication device" is a network-connected server used to send and receive data.

[0106] "Artificial intelligence" is a technology in which computers mimic human intellectual tasks and perform functions such as natural language processing.

[0107] An "identification image" is a visual element provided in the form of a code or tag for digitizing information and extracting its content.

[0108] The terminal is a device that stores the user's selected language settings and translates information based on them. When the terminal accesses an information source such as a web page or application, it first detects the display language and acquires that information as an identification image. Based on the acquired information, it generates a translation request and sends it to the communication device. This communication device functions as a server and, based on the received translation request, uses a generating AI to translate the original data. Specifically, it performs the translation using a generated artificial intelligence language model—for example, a modern neural network model. The translated content is sent back to the terminal, which displays it in the corresponding location within the user interface.

[0109] This system allows users to quickly grasp information and take appropriate action without experiencing stress, even in different language environments. For example, one scenario could involve scanning a QR code (registered trademark) of a product sold in Japan, translating the product details into English, and proceeding with the purchase. In this case, the prompt would be something like, "Please translate into English."

[0110] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0111] Step 1:

[0112] The device saves the language settings specified by the user. During this process, it retrieves the desired language data entered by the user through the device's language settings interface and writes it to the device's memory. This saved data is then used in subsequent processing.

[0113] Step 2:

[0114] The terminal detects the display language of the information source being accessed. It analyzes text information obtained from web pages and applications and applies a language recognition algorithm to determine the language of the input data. This process is generally performed using a natural language processing library, and the detected language code is generated as output.

[0115] Step 3:

[0116] The terminal creates a translation request based on the detected language and the language specified by the user and sends it to the communication device. In this process, it receives the detected language, target language, and text information to be translated as input, and constructs a translation request data package that includes a prompt. As output, data in a format suitable for the communication protocol is generated.

[0117] Step 4:

[0118] The server processes translation requests received from terminals and performs data translation using a generative AI model. Following the translation instructions in the request, the original text is input into the language model, and the AI ​​outputs the text result converted into the target language specified by the AI. This process utilizes advanced neural language models.

[0119] Step 5:

[0120] The server sends the translation results to the terminal. The generated translated text is formatted into a data packet and sent back to the terminal via the communication line. The output data is transmitted in an appropriate format considering the communication conditions.

[0121] Step 6:

[0122] The device reflects the received translation data in the interface and displays it to the user. The acquired translated text is applied to UI elements and output in a visually verifiable format. The user can then decide on their actions based on this information.

[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0124] This invention is a multilingual system combining a terminal, a server, a generative AI, and an emotion engine. This system enables users to communicate smoothly across language barriers. The user begins using the terminal by launching an application or website and selecting their desired language. The terminal saves this language setting and automatically detects the language of the displayed content.

[0125] The device then generates a translation request and sends it to the server. This request includes the original display language, the user's selected language, and the specific text information to be translated. The server processes the received request and forwards it to the generation AI. This AI uses a highly accurate neural model to translate the text into the user's desired language.

[0126] Furthermore, the addition of an emotion engine allows the system to analyze user voice input and facial expression data to recognize the user's emotions. The server then uses the results of emotion recognition to apply the most appropriate expression to the situation through a generation AI, adjusting the translation result accordingly.

[0127] The translated and emotion-adjusted text is returned from the server to the device and displayed in the user interface in the most optimal format. As a concrete example, consider a scenario where a user uses a Japanese tourist information app and smiles with satisfaction after reading the information. The emotion engine detects positive emotions from the user's facial expressions and provides a soft, friendly translation that matches their intent. This allows the user to use the app without stress and have a better experience.

[0128] This system not only translates language but also presents content that reflects the user's emotional state, making it possible to provide information more appropriately and effectively to people from diverse cultural backgrounds.

[0129] The following describes the processing flow.

[0130] Step 1:

[0131] The user launches the application on their device and selects their desired language on the language settings screen. The device saves this language setting to the user profile.

[0132] Step 2:

[0133] The device activates a translation agent and automatically detects the display language of the website or application the user is trying to view. This detection is performed by analyzing the text and metadata on the page.

[0134] Step 3:

[0135] The device extracts the necessary text based on the detected display language and the user's selected language, and generates a translation request for the server. This request includes the specific text information that needs to be translated.

[0136] Step 4:

[0137] The server receives the translation request and sends a request to the generative AI to perform the translation. The generative AI uses a multilingual translation model to perform an accurate translation from the original language to the user's selected language.

[0138] Step 5:

[0139] The device also acquires data on the user's facial expressions and voice and sends it to the emotion engine. The emotion engine analyzes this data to identify the user's emotional state.

[0140] Step 6:

[0141] The server optimizes the translation results created by the generative AI based on emotional data from the emotion engine. It applies translation expressions that match the emotion and adjusts them so that the message is conveyed to the user in the most appropriate way.

[0142] Step 7:

[0143] The server sends the adjusted translation results to the terminal. The terminal then reflects the translated text in the user interface of the application or website and displays it to the user.

[0144] Step 8:

[0145] Through translated content, users can receive flexible information tailored to their chosen language and sentiment, enabling them to perform meaningful actions.

[0146] (Example 2)

[0147] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0148] In addition to facilitating communication between multiple languages, it is necessary to address the challenge of providing translations that take into account the user's emotional state. Conventional systems are limited to language conversion and have difficulty providing flexible translations that respond to the user's emotions and circumstances.

[0149] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0150] In this invention, the server includes means for converting character information with high accuracy based on received conversion requests, means for generating translation results using generative artificial intelligence, and means for recognizing the user's emotions using an emotion analysis device and adjusting the translation results accordingly. This makes it possible to provide flexible and natural translation results that reflect the user's emotions.

[0151] A "terminal" refers to an information processing device, a device that the user directly operates.

[0152] An "information processing device" is a part of a computer system and is a device used to process digital content.

[0153] "Digital content" is a general term for information and media that are expressed and stored electronically.

[0154] "Display language" refers to the language used by information processing devices and digital content.

[0155] A "conversion request" is a command generated and sent to a communication device for the purpose of converting languages.

[0156] A "communication device" is a part of the hardware or software used to send and receive data and information between terminals and other devices.

[0157] "Textual information" refers to linguistic data expressed in text format.

[0158] "Generative artificial intelligence" refers to artificial systems that use technologies such as machine learning and neural networks to mimic human intellectual activity.

[0159] An "emotion analysis device" is a device that analyzes a user's emotional state and provides appropriate feedback.

[0160] A "translation result" is a text in a new language format generated based on a conversion request.

[0161] A "user interface" refers to the screens and operating methods that a user uses to interact with a device or application.

[0162] This invention is a system that combines a terminal operated directly by the user, an information processing device for processing digital content, and a communication device for sending and receiving data. The user can start the system by using a designated application or website on the terminal and selecting their desired language. The terminal stores this language setting internally and uses that information in subsequent processes.

[0163] The terminal utilizes natural language processing technology to detect the language of the digital content being displayed. This identifies the basic linguistic characteristics of the content and generates a translation request. The generated translation request includes the detected language, the user-selected language, and the character information that needs to be translated. The translation request is sent to the server via a communication device.

[0164] The server uses a highly trained generative artificial intelligence model to translate incoming text information with high accuracy. During this process, an emotion analyzer analyzes the user's voice and facial expressions to identify their emotional state. Based on this emotion data, the server adjusts the translation results, going beyond simple language conversion to provide natural and appropriate expressions that reflect the user's emotions.

[0165] As a concrete example, consider a scenario where a user is using a Japanese tourist information app and wants to understand the guide information in English. If the user is smiling and looking at the screen, the emotion analyzer will detect his positive emotions, and the server can return a friendly translation accordingly.

[0166] An example of a prompt for a generative AI model would be text like, "Translate this Japanese tourist guide into English. The user is smiling." Based on this prompt, the system provides the best possible translation in the specified language.

[0167] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0168] Step 1:

[0169] The user launches an application or website on their device and selects their desired language. This input prompts the device to save the user's language settings to its data storage. This saved information is then used for subsequent language processing.

[0170] Step 2:

[0171] The terminal detects the language of the digital content being displayed. The input is content data, and natural language processing techniques are applied to analyze its linguistic characteristics. The output is information about the detected display language, which forms the basis for creating translation requests.

[0172] Step 3:

[0173] The terminal generates a translation request based on the detected display language and the language selected by the user. The input is the language detection result and the user's selected language, and the output is request data, which includes the character information to be translated. This request is sent to the server via a communication device.

[0174] Step 4:

[0175] The server uses artificial intelligence to translate text information based on the received translation request. The input is text information received from the terminal, and the AI's neural network is used to perform language conversion. The output is translated text data.

[0176] Step 5:

[0177] When a user provides voice input or displays facial expressions, the terminal processes this data using an emotion analysis device. The input consists of the user's real-time voice and visual data, and an emotion recognition algorithm is applied to identify the user's emotional state. The output is the analyzed emotion information.

[0178] Step 6:

[0179] The server adjusts the translation results based on the sentiment analysis. The input consists of the AI-generated translation results and sentiment data, and the tone and expression are adjusted to achieve natural dialogue. The output is the optimally translated text that corresponds to the sentiment.

[0180] Step 7:

[0181] The server returns the final translated text to the terminal. The output text is transmitted to the terminal via a communication device and displayed on the user interface. This allows the user to visually confirm the translation results in an appropriate format.

[0182] (Application Example 2)

[0183] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0184] In multilingual communication, it is necessary not only to convert text into text, but also to accurately convey the speaker's emotions and intentions. However, conventional systems have limitations in language translation accuracy and fail to provide appropriate translations that take into account the user's emotions. In particular, in fields such as tourist information, there is a need to provide accurate and user-friendly information that aligns with the user's feelings.

[0185] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0186] In this invention, the server includes means for converting text, means for recognizing and adjusting the user's emotions from their facial expressions and voice, and means for transmitting the converted information to a terminal. This makes it possible to provide translation results that are adapted to the user's emotions.

[0187] A "terminal" is an electronic device used by a user, and is a device for inputting and displaying information.

[0188] "Language settings" refer to settings that allow users to specify their preferred language.

[0189] An "information page" is a general term for informational content displayed on online web pages and applications.

[0190] A "symbol" is an element that represents information, such as letters or words, in natural language.

[0191] A "translation request" is request information generated to convert text into another language.

[0192] An "information processing device" is a system that processes received data and sends a response to a terminal.

[0193] "Textual information" refers to data in text format, which is information written in human language.

[0194] An "emotion analysis device" is a system for analyzing a user's emotions, and it recognizes emotions using voice and facial expression data.

[0195] A "dialogue screen" is the screen portion that displays information via the user interface.

[0196] To implement this application, a system is built using a server as the information processing device and the user's terminal as the client. The user uses a dedicated application installed on the terminal to set the desired language and then begins accessing the information page. The terminal detects the displayed symbols and sends a translation request to the server based on them.

[0197] To process incoming translation requests, the server first analyzes the text information and performs the appropriate conversion. The software used includes extracting text from images using the Google® Cloud Vision API and translating the language using the Azure® Translator API. Furthermore, an emotion analyzer recognizes emotions from the user's facial expressions and voice. In this process, the Face API and Emotion API from Microsoft® Azure Cognitive Services are used to evaluate the user's emotional state.

[0198] Based on recognized emotions, a generative AI model is used to optimize the translation results. This generates flexible translations that are sensitive to the user's emotions. The optimized translation, using the generated prompt text, is sent to the device and displayed through the user interface.

[0199] As a concrete example, when a user takes a picture of a sign at a tourist attraction with their camera, the information processing device analyzes the image and translates it into the user's preferred language. At the same time, if it determines that the user is smiling and showing interest, it generates a translation that gently conveys detailed background information.

[0200] An example of a prompt for a generative AI model would be: "Translate the following Japanese text into English and add a detailed explanation if the user is interested: Japanese text."

[0201] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0202] Step 1:

[0203] The user launches the application on their device and sets their desired language. The input is the user's language selection, and the output is the set language information saved on the device. This language setting is used in the subsequent translation process.

[0204] Step 2:

[0205] The device accesses an information page and detects the displayed symbols within it. The input is the information page accessed by the user, and the output is the detected symbol data. This is done by extracting text from the image using the Google Cloud Vision API.

[0206] Step 3:

[0207] The terminal generates a translation request based on the detected symbol data and stored language settings, and sends it to the server. The input is the symbol data and language settings, and the output is the translation request data sent to the server. This prepares the server to begin the translation process.

[0208] Step 4:

[0209] The server analyzes the character information based on the received translation request and uses the Azure Translator API to translate the symbolic data into the user's desired language. The input is the translation request data, and the output is the translated text data. Here, data transformation for translation takes place.

[0210] Step 5:

[0211] The server uses an emotion analysis device to analyze the user's facial expressions and voice in order to understand the user's emotional state. The input is facial and voice data sent from the terminal, and the output is a recognition of the user's emotional state. Emotion analysis is performed using Microsoft Azure Cognitive Services' Face API and Emotion API.

[0212] Step 6:

[0213] The server optimizes the translation results using a generative AI model based on the recognized emotional state. The input is the emotional state and the initial translated text; the output is a flexible translation that reflects the emotion. The generative AI model makes appropriate adjustments using prompt sentences.

[0214] Step 7:

[0215] The server sends optimized translation results to the terminal, which displays the translation results through the user interface. The input is the optimized translation data received from the server, and the output is the translation result displayed to the user. This allows the user to receive translation results that are adapted to their emotions.

[0216] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0217] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0218] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0219] [Second Embodiment]

[0220] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0221] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0222] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0223] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0224] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0225] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0226] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0227] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0228] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0229] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0230] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0231] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0232] This invention is a multilingual system utilizing a terminal, a server, and a generative AI. The terminal used by the user simultaneously launches its own translation agent when an application or website is started. This translation agent records the language selected by the user on the terminal as initial setting data and automatically detects the display language of the content the terminal is accessing.

[0233] Based on the detected display language, the terminal generates and sends a translation request to the server. This translation request includes the original language, the target language (the language selected by the user), and the text information that needs to be translated.

[0234] The server processes the received request and forwards it to the generative AI. The generative AI uses advanced language models, such as neural network models, to translate the received text into the user's selected language. The translated text information is then sent to the terminal by the server as the translation result.

[0235] The device translates and displays the relevant parts of the application or website's user interface based on the translation results received from the server. This allows the user to view the information in their chosen language.

[0236] As a concrete example, consider a scenario where a user who only speaks English uses a Japanese public transportation timetable app. The user launches the app and sets English as the preferred language. The device's translation agent detects the timetable information provided in Japanese and sends a request to the server to translate it into English. The server uses generative AI to translate the information, and the translated timetable is displayed on the device. Based on this, the user can plan their travel schedule and enjoy a comfortable sightseeing experience.

[0237] The following describes the processing flow.

[0238] Step 1:

[0239] The user launches the application on their device and sets their preferred language. This setting information is saved on the device and used in subsequent processes.

[0240] Step 2:

[0241] The device activates a translation agent and detects the display language of the content the user is viewing. This detection is performed by analyzing the text on the page.

[0242] Step 3:

[0243] The device creates a translation request based on the detected display language and the language selected by the user, and sends it to the server. This request also includes the specific text that needs to be translated.

[0244] Step 4:

[0245] The server processes the received translation request and passes it to the generating AI. The AI ​​model translates the requested text into the user's selected language.

[0246] Step 5:

[0247] The server sends the translated result to the terminal. This result is data containing the translated text.

[0248] Step 6:

[0249] The device translates the user interface of an application or website based on the received translation results and displays it to the user.

[0250] Step 7:

[0251] Users can perform the necessary actions using translated content, thereby enabling them to accurately understand and utilize the information.

[0252] (Example 1)

[0253] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0254] When accessing information in different languages, users often face challenges in understanding the information, particularly in the interface, which can make smooth operation difficult. This means that language barriers are a significant obstacle to information utilization and international communication.

[0255] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0256] In this invention, the server includes means for converting character data based on a received translation request, means for transmitting the conversion result to a terminal, and means for receiving translation requests from the terminal. This enables smooth translation and display of information in different languages.

[0257] A "terminal" is an electronic device that users directly operate to access information.

[0258] "Language settings" is a component that allows users to specify their preferred language for displaying and interacting with information.

[0259] "Display language" refers to the language used when information or an interface is presented visually to the user.

[0260] "Means of recognition" refers to a function that allows the terminal to automatically identify the language in which information is displayed.

[0261] A "translation request" is a request sent from a terminal to a server that includes the source language and the target language of the information.

[0262] A "device for aggregating information" is a central computer that performs data conversion and management based on requests from terminals.

[0263] "Means of converting text data" refers to the process of automatically translating information into another language.

[0264] "Conversion result" refers to the output after converting character data to another language.

[0265] "User interface" refers to the screen that the user directly interacts with, where information is presented through the terminal.

[0266] "Generative artificial intelligence" is a technology that uses neural network models to automatically translate text into other languages.

[0267] This invention is a system that facilitates multilingual information access and includes a terminal for user access, a server for aggregating information, and a generative AI with advanced translation capabilities.

[0268] The user launches a specific source of information (e.g., a webpage or application) on their device. This device has a feature that allows the user to specify their preferred display language as the default setting. The device automatically parses HTML tags and metadata to recognize the language of the displayed information.

[0269] The device generates a translation request to a server that aggregates information, based on its recognized language settings. This request includes the original language, the user's desired target language, and the text to be translated. The server receives this request and sends it as a prompt to the artificial intelligence that generates the information.

[0270] The generative AI uses a neural network model to translate text into the desired language. This AI considers context and cultural nuances to produce natural-sounding translations. The translated text is then sent to the terminal via a server.

[0271] The terminal displays the received conversion result in the user's preferred language on the interface they interact with. This allows the user to understand and interact with information in their chosen language.

[0272] As a concrete example, consider a scenario where a user who only understands English accesses a tourist information webpage provided in Japanese. In this case, the user opens the webpage and specifies English as the desired display language. The device recognizes the Japanese content and sends a request to the server to translate it into English. The server uses generative AI to translate the information, and the English content is displayed on the device. This allows the user to use the information without being aware of the language barrier.

[0273] An example of a prompt message would be the text format: "Please translate the following Japanese sentence into English: 'Hello'".

[0274] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0275] Step 1:

[0276] The user launches a specific application or website using the device. The translation agent on the device records the language the user initially selected. Based on this input, the device retains the user's preferred language as output.

[0277] Step 2:

[0278] The device automatically detects the display language of the content being accessed. This process involves analyzing HTML tags and metadata, for example, by checking the "lang" attribute to identify the language. The detected language is then used as input data for the translation process.

[0279] Step 3:

[0280] The device generates a translation request based on the detected display language and the user's language setting, and sends it to the server. This translation request includes the source language, the target language, and the text to be translated. This is then sent to the server as output data for the translation process.

[0281] Step 4:

[0282] The server analyzes the translation request received from the terminal and forwards the request to the generating AI model along with the generated prompt. The server's input is the translation request, and its output is the AI ​​request, which includes the prompt.

[0283] Step 5:

[0284] The generative AI performs data processing for natural translation based on the received prompt. The AI analyzes the context and nuances using a neural network and outputs the translated text. The output is the translated text in the language selected by the user.

[0285] Step 6:

[0286] The server receives the translated text from the generative AI and sends it to the user's terminal. The input to the server is the translation result of the AI, and the output is the conversion data to the terminal.

[0287] Step 7:

[0288] Based on the received translated text, the terminal translates the necessary parts in the user interface. As a result, the user can check the information in their set language. The input is the translation result from the server, and the output is the translated user interface.

[0289] (Application Example 1)

[0290] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0291] In a multilingual environment, when a user purchases a product or makes a payment, it is an issue to reduce the difficulties caused by the language barrier and enable the user to conduct business transactions with confidence. In particular, when the visited country or region uses a different language, it is required to quickly and accurately translate the information into the user's mother tongue and respond smoothly.

[0292] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0293] In this invention, the server includes a device that stores the user's selected language settings on the terminal, a device that acquires information presented using an identification image on the terminal and generates a translation request, and a device that performs the translation process using the generated artificial intelligence. This enables users to quickly verify information in their native language and conduct transactions reliably, even in different language environments.

[0294] A "terminal" is an electronic device that a user directly operates to display information.

[0295] A "user" is defined as the entity that uses the system to translate or verify information.

[0296] "Language settings" means that the user specifies their native language or preferred language.

[0297] A "device" is a component of hardware or software designed to perform a specific function.

[0298] "Information sources" refer to websites and applications accessible through a device.

[0299] "Detection" is the act of analyzing information in order to identify a specified language.

[0300] A "translation request" is the process of requesting that information be converted from the original language to the user's desired language.

[0301] A "communication device" is a network-connected server used to send and receive data.

[0302] "Artificial intelligence" is a technology in which computers mimic human intellectual tasks and perform functions such as natural language processing.

[0303] An "identification image" is a visual element provided in the form of a code or tag for digitizing information and extracting its content.

[0304] The terminal is a device that stores the language settings selected by the user and performs information translation based on them. When the terminal accesses an information source such as a web page or an application, it first detects the display language and acquires the information as an identification image. Based on the acquired information, a translation request is generated and sent to the communication device. This communication device functions as a server and translates the original data using a generation AI based on the received translation request. Specifically, a language model, which is an artificial intelligence generated, such as the latest neural network model, is used to perform the translation. The translated content is sent back to the terminal, and the terminal displays it at the corresponding location within the user interface.

[0305] With this system, users can quickly understand information and take appropriate actions without feeling stressed in different language environments. For example, a scenario where a QR code of a product sold in Japan is scanned, the product details are translated into English, and the purchase process is advanced can be considered. As an example of the prompt text used in this case, the input content is specified in the form of "Please translate it into English".

[0306] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0307] Step 1:

[0308] The terminal stores the language settings specified by the user. At this time, the data of the desired language input by the user through the language setting interface of the terminal is acquired and written into the memory of the device. The stored data is used in subsequent processing.

[0309] Step 2:

[0310] The terminal detects the display language of the information source being accessed. It analyzes text information obtained from web pages and applications and applies a language recognition algorithm to determine the language of the input data. This process is generally performed using a natural language processing library, and the detected language code is generated as output.

[0311] Step 3:

[0312] The terminal creates a translation request based on the detected language and the language specified by the user and sends it to the communication device. In this process, it receives the detected language, target language, and text information to be translated as input, and constructs a translation request data package that includes a prompt. As output, data in a format suitable for the communication protocol is generated.

[0313] Step 4:

[0314] The server processes translation requests received from terminals and performs data translation using a generative AI model. Following the translation instructions in the request, the original text is input into the language model, and the AI ​​outputs the text result converted into the target language specified by the AI. This process utilizes advanced neural language models.

[0315] Step 5:

[0316] The server sends the translation results to the terminal. The generated translated text is formatted into a data packet and sent back to the terminal via the communication line. The output data is transmitted in an appropriate format considering the communication conditions.

[0317] Step 6:

[0318] The device reflects the received translation data in the interface and displays it to the user. The acquired translated text is applied to UI elements and output in a visually verifiable format. The user can then decide on their actions based on this information.

[0319] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0320] This invention is a multilingual system combining a terminal, a server, a generative AI, and an emotion engine. This system enables users to communicate smoothly across language barriers. The user begins using the terminal by launching an application or website and selecting their desired language. The terminal saves this language setting and automatically detects the language of the displayed content.

[0321] The device then generates a translation request and sends it to the server. This request includes the original display language, the user's selected language, and the specific text information to be translated. The server processes the received request and forwards it to the generation AI. This AI uses a highly accurate neural model to translate the text into the user's desired language.

[0322] Furthermore, the addition of an emotion engine allows the system to analyze user voice input and facial expression data to recognize the user's emotions. The server then uses the results of emotion recognition to apply the most appropriate expression to the situation through a generation AI, adjusting the translation result accordingly.

[0323] The translated and emotion-adjusted text is returned from the server to the device and displayed in the user interface in the most optimal format. As a concrete example, consider a scenario where a user uses a Japanese tourist information app and smiles with satisfaction after reading the information. The emotion engine detects positive emotions from the user's facial expressions and provides a soft, friendly translation that matches their intent. This allows the user to use the app without stress and have a better experience.

[0324] This system not only translates language but also presents content that reflects the user's emotional state, making it possible to provide information more appropriately and effectively to people from diverse cultural backgrounds.

[0325] The following describes the processing flow.

[0326] Step 1:

[0327] The user launches the application on their device and selects their desired language on the language settings screen. The device saves this language setting to the user profile.

[0328] Step 2:

[0329] The device activates a translation agent and automatically detects the display language of the website or application the user is trying to view. This detection is performed by analyzing the text and metadata on the page.

[0330] Step 3:

[0331] The device extracts the necessary text based on the detected display language and the user's selected language, and generates a translation request for the server. This request includes the specific text information that needs to be translated.

[0332] Step 4:

[0333] The server receives the translation request and sends a request to the generative AI to perform the translation. The generative AI uses a multilingual translation model to perform an accurate translation from the original language to the user's selected language.

[0334] Step 5:

[0335] The device also acquires data on the user's facial expressions and voice and sends it to the emotion engine. The emotion engine analyzes this data to identify the user's emotional state.

[0336] Step 6:

[0337] The server optimizes the translation results created by the generative AI based on emotional data from the emotion engine. It applies translation expressions that match the emotion and adjusts them so that the message is conveyed to the user in the most appropriate way.

[0338] Step 7:

[0339] The server sends the adjusted translation results to the terminal. The terminal then reflects the translated text in the user interface of the application or website and displays it to the user.

[0340] Step 8:

[0341] Through translated content, users can receive flexible information tailored to their chosen language and sentiment, enabling them to perform meaningful actions.

[0342] (Example 2)

[0343] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0344] In addition to facilitating communication between multiple languages, it is necessary to address the challenge of providing translations that take into account the user's emotional state. Conventional systems are limited to language conversion and have difficulty providing flexible translations that respond to the user's emotions and circumstances.

[0345] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0346] In this invention, the server includes means for converting character information with high accuracy based on received conversion requests, means for generating translation results using generative artificial intelligence, and means for recognizing the user's emotions using an emotion analysis device and adjusting the translation results accordingly. This makes it possible to provide flexible and natural translation results that reflect the user's emotions.

[0347] A "terminal" refers to an information processing device, a device that the user directly operates.

[0348] An "information processing device" is a part of a computer system and is a device used to process digital content.

[0349] "Digital content" is a general term for information and media that are expressed and stored electronically.

[0350] "Display language" refers to the language used by information processing devices and digital content.

[0351] A "conversion request" is a command generated and sent to a communication device for the purpose of converting languages.

[0352] A "communication device" is a part of the hardware or software used to send and receive data and information between terminals and other devices.

[0353] "Textual information" refers to linguistic data expressed in text format.

[0354] "Generative artificial intelligence" refers to artificial systems that use technologies such as machine learning and neural networks to mimic human intellectual activity.

[0355] An "emotion analysis device" is a device that analyzes a user's emotional state and provides appropriate feedback.

[0356] A "translation result" is a text in a new language format generated based on a conversion request.

[0357] A "user interface" refers to the screens and operating methods that a user uses to interact with a device or application.

[0358] This invention is a system that combines a terminal operated directly by the user, an information processing device for processing digital content, and a communication device for sending and receiving data. The user can start the system by using a designated application or website on the terminal and selecting their desired language. The terminal stores this language setting internally and uses that information in subsequent processes.

[0359] The terminal utilizes natural language processing technology to detect the language of the digital content being displayed. This identifies the basic linguistic characteristics of the content and generates a translation request. The generated translation request includes the detected language, the user-selected language, and the character information that needs to be translated. The translation request is sent to the server via a communication device.

[0360] The server uses a highly trained generative artificial intelligence model to translate incoming text information with high accuracy. During this process, an emotion analyzer analyzes the user's voice and facial expressions to identify their emotional state. Based on this emotion data, the server adjusts the translation results, going beyond simple language conversion to provide natural and appropriate expressions that reflect the user's emotions.

[0361] As a concrete example, consider a scenario where a user is using a Japanese tourist information app and wants to understand the guide information in English. If the user is smiling and looking at the screen, the emotion analyzer will detect his positive emotions, and the server can return a friendly translation accordingly.

[0362] An example of a prompt for a generative AI model would be text like, "Translate this Japanese tourist guide into English. The user is smiling." Based on this prompt, the system provides the best possible translation in the specified language.

[0363] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0364] Step 1:

[0365] The user launches an application or website on their device and selects their desired language. This input prompts the device to save the user's language settings to its data storage. This saved information is then used for subsequent language processing.

[0366] Step 2:

[0367] The terminal detects the language of the digital content being displayed. The input is content data, and natural language processing techniques are applied to analyze its linguistic characteristics. The output is information about the detected display language, which forms the basis for creating translation requests.

[0368] Step 3:

[0369] The terminal generates a translation request based on the detected display language and the language selected by the user. The input is the language detection result and the user's selected language, and the output is request data, which includes the character information to be translated. This request is sent to the server via a communication device.

[0370] Step 4:

[0371] The server uses artificial intelligence to translate text information based on the received translation request. The input is text information received from the terminal, and the AI's neural network is used to perform language conversion. The output is translated text data.

[0372] Step 5:

[0373] When a user provides voice input or displays facial expressions, the terminal processes this data using an emotion analysis device. The input consists of the user's real-time voice and visual data, and an emotion recognition algorithm is applied to identify the user's emotional state. The output is the analyzed emotion information.

[0374] Step 6:

[0375] The server adjusts the translation results based on the sentiment analysis. The input consists of the AI-generated translation results and sentiment data, and the tone and expression are adjusted to achieve natural dialogue. The output is the optimally translated text that corresponds to the sentiment.

[0376] Step 7:

[0377] The server returns the final translated text to the terminal. The output text is transmitted to the terminal via a communication device and displayed on the user interface. This allows the user to visually confirm the translation results in an appropriate format.

[0378] (Application Example 2)

[0379] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0380] In multilingual communication, it is necessary not only to convert text into text, but also to accurately convey the speaker's emotions and intentions. However, conventional systems have limitations in language translation accuracy and fail to provide appropriate translations that take into account the user's emotions. In particular, in fields such as tourist information, there is a need to provide accurate and user-friendly information that aligns with the user's feelings.

[0381] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0382] In this invention, the server includes means for converting text, means for recognizing and adjusting the user's emotions from their facial expressions and voice, and means for transmitting the converted information to a terminal. This makes it possible to provide translation results that are adapted to the user's emotions.

[0383] A "terminal" is an electronic device used by a user, and is a device for inputting and displaying information.

[0384] "Language settings" refer to settings that allow users to specify their preferred language.

[0385] An "information page" is a general term for informational content displayed on online web pages and applications.

[0386] A "symbol" is an element that represents information, such as letters or words, in natural language.

[0387] A "translation request" is request information generated to convert text into another language.

[0388] An "information processing device" is a system that processes received data and sends a response to a terminal.

[0389] "Textual information" refers to data in text format, which is information written in human language.

[0390] An "emotion analysis device" is a system for analyzing a user's emotions, and it recognizes emotions using voice and facial expression data.

[0391] A "dialogue screen" is the screen portion that displays information via the user interface.

[0392] To implement this application, a system is built using a server as the information processing device and the user's terminal as the client. The user uses a dedicated application installed on the terminal to set the desired language and then begins accessing the information page. The terminal detects the displayed symbols and sends a translation request to the server based on them.

[0393] To process incoming translation requests, the server first analyzes the text information and performs the appropriate conversion. The software used includes extracting text from images using the Google Cloud Vision API and translating the language using the Azure Translator API. Furthermore, an emotion analyzer recognizes emotions from the user's facial expressions and voice. In this process, the Face API and Emotion API from Microsoft Azure Cognitive Services are used to evaluate the user's emotional state.

[0394] Based on recognized emotions, a generative AI model is used to optimize the translation results. This generates flexible translations that are sensitive to the user's emotions. The optimized translation, using the generated prompt text, is sent to the device and displayed through the user interface.

[0395] As a concrete example, when a user takes a picture of a sign at a tourist attraction with their camera, the information processing device analyzes the image and translates it into the user's preferred language. At the same time, if it determines that the user is smiling and showing interest, it generates a translation that gently conveys detailed background information.

[0396] An example of a prompt for a generative AI model would be: "Translate the following Japanese text into English and add a detailed explanation if the user is interested: Japanese text."

[0397] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0398] Step 1:

[0399] The user launches the application on their device and sets their desired language. The input is the user's language selection, and the output is the set language information saved on the device. This language setting is used in the subsequent translation process.

[0400] Step 2:

[0401] The device accesses an information page and detects the displayed symbols within it. The input is the information page accessed by the user, and the output is the detected symbol data. This is done by extracting text from the image using the Google Cloud Vision API.

[0402] Step 3:

[0403] The terminal generates a translation request based on the detected symbol data and stored language settings, and sends it to the server. The input is the symbol data and language settings, and the output is the translation request data sent to the server. This prepares the server to begin the translation process.

[0404] Step 4:

[0405] The server analyzes the character information based on the received translation request and uses the Azure Translator API to translate the symbolic data into the user's desired language. The input is the translation request data, and the output is the translated text data. Here, data transformation for translation takes place.

[0406] Step 5:

[0407] The server uses an emotion analysis device to analyze the user's facial expressions and voice in order to understand the user's emotional state. The input is facial and voice data sent from the terminal, and the output is a recognition of the user's emotional state. Emotion analysis is performed using Microsoft Azure Cognitive Services' Face API and Emotion API.

[0408] Step 6:

[0409] The server optimizes the translation results using a generative AI model based on the recognized emotional state. The input is the emotional state and the initial translated text; the output is a flexible translation that reflects the emotion. The generative AI model makes appropriate adjustments using prompt sentences.

[0410] Step 7:

[0411] The server sends optimized translation results to the terminal, which displays the translation results through the user interface. The input is the optimized translation data received from the server, and the output is the translation result displayed to the user. This allows the user to receive translation results that are adapted to their emotions.

[0412] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0413] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0414] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0415] [Third Embodiment]

[0416] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0417] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0418] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0419] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0420] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0421] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0422] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0423] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0424] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0425] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0426] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0427] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0428] This invention is a multilingual system utilizing a terminal, a server, and a generative AI. The terminal used by the user simultaneously launches its own translation agent when an application or website is started. This translation agent records the language selected by the user on the terminal as initial setting data and automatically detects the display language of the content the terminal is accessing.

[0429] Based on the detected display language, the terminal generates and sends a translation request to the server. This translation request includes the original language, the target language (the language selected by the user), and the text information that needs to be translated.

[0430] The server processes the received request and forwards it to the generative AI. The generative AI uses advanced language models, such as neural network models, to translate the received text into the user's selected language. The translated text information is then sent to the terminal by the server as the translation result.

[0431] The device translates and displays the relevant parts of the application or website's user interface based on the translation results received from the server. This allows the user to view the information in their chosen language.

[0432] As a concrete example, consider a scenario where a user who only speaks English uses a Japanese public transportation timetable app. The user launches the app and sets English as the preferred language. The device's translation agent detects the timetable information provided in Japanese and sends a request to the server to translate it into English. The server uses generative AI to translate the information, and the translated timetable is displayed on the device. Based on this, the user can plan their travel schedule and enjoy a comfortable sightseeing experience.

[0433] The following describes the processing flow.

[0434] Step 1:

[0435] The user launches the application on their device and sets their preferred language. This setting information is saved on the device and used in subsequent processes.

[0436] Step 2:

[0437] The device activates a translation agent and detects the display language of the content the user is viewing. This detection is performed by analyzing the text on the page.

[0438] Step 3:

[0439] The device creates a translation request based on the detected display language and the language selected by the user, and sends it to the server. This request also includes the specific text that needs to be translated.

[0440] Step 4:

[0441] The server processes the received translation request and passes it to the generating AI. The AI ​​model translates the requested text into the user's selected language.

[0442] Step 5:

[0443] The server sends the translated result to the terminal. This result is data containing the translated text.

[0444] Step 6:

[0445] The device translates the user interface of an application or website based on the received translation results and displays it to the user.

[0446] Step 7:

[0447] Users can perform the necessary actions using translated content, thereby enabling them to accurately understand and utilize the information.

[0448] (Example 1)

[0449] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0450] When accessing information in different languages, users often face challenges in understanding the information, particularly in the interface, which can make smooth operation difficult. This means that language barriers are a significant obstacle to information utilization and international communication.

[0451] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0452] In this invention, the server includes means for converting character data based on a received translation request, means for transmitting the conversion result to a terminal, and means for receiving translation requests from the terminal. This enables smooth translation and display of information in different languages.

[0453] A "terminal" is an electronic device that users directly operate to access information.

[0454] "Language settings" is a component that allows users to specify their preferred language for displaying and interacting with information.

[0455] "Display language" refers to the language used when information or an interface is presented visually to the user.

[0456] "Means of recognition" refers to a function that allows the terminal to automatically identify the language in which information is displayed.

[0457] A "translation request" is a request sent from a terminal to a server that includes the source language and the target language of the information.

[0458] A "device for aggregating information" is a central computer that performs data conversion and management based on requests from terminals.

[0459] "Means of converting text data" refers to the process of automatically translating information into another language.

[0460] "Conversion result" refers to the output after converting character data to another language.

[0461] "User interface" refers to the screen that the user directly interacts with, where information is presented through the terminal.

[0462] "Generative artificial intelligence" is a technology that uses neural network models to automatically translate text into other languages.

[0463] This invention is a system that facilitates multilingual information access and includes a terminal for user access, a server for aggregating information, and a generative AI with advanced translation capabilities.

[0464] The user launches a specific source of information (e.g., a webpage or application) on their device. This device has a feature that allows the user to specify their preferred display language as the default setting. The device automatically parses HTML tags and metadata to recognize the language of the displayed information.

[0465] The device generates a translation request to a server that aggregates information, based on its recognized language settings. This request includes the original language, the user's desired target language, and the text to be translated. The server receives this request and sends it as a prompt to the artificial intelligence that generates the information.

[0466] The generative AI uses a neural network model to translate text into the desired language. This AI considers context and cultural nuances to produce natural-sounding translations. The translated text is then sent to the terminal via a server.

[0467] The terminal displays the received conversion result in the user's preferred language on the interface they interact with. This allows the user to understand and interact with information in their chosen language.

[0468] As a concrete example, consider a scenario where a user who only understands English accesses a tourist information webpage provided in Japanese. In this case, the user opens the webpage and specifies English as the desired display language. The device recognizes the Japanese content and sends a request to the server to translate it into English. The server uses generative AI to translate the information, and the English content is displayed on the device. This allows the user to use the information without being aware of the language barrier.

[0469] An example of a prompt message would be the text format: "Please translate the following Japanese sentence into English: 'Hello'".

[0470] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0471] Step 1:

[0472] The user launches a specific application or website using the device. The translation agent on the device records the language the user initially selected. Based on this input, the device retains the user's preferred language as output.

[0473] Step 2:

[0474] The device automatically detects the display language of the content being accessed. This process involves analyzing HTML tags and metadata, for example, by checking the "lang" attribute to identify the language. The detected language is then used as input data for the translation process.

[0475] Step 3:

[0476] The device generates a translation request based on the detected display language and the user's language setting, and sends it to the server. This translation request includes the source language, the target language, and the text to be translated. This is then sent to the server as output data for the translation process.

[0477] Step 4:

[0478] The server analyzes the translation request received from the terminal and forwards the request to the generating AI model along with the generated prompt. The server's input is the translation request, and its output is the AI ​​request, which includes the prompt.

[0479] Step 5:

[0480] The generative AI processes data based on the received prompt to produce a natural translation. The AI ​​uses a neural network to analyze context and nuances, and outputs the translated text. The output is the translated text in the language selected by the user.

[0481] Step 6:

[0482] The server receives the translated text from the generating AI and sends it to the user's terminal. The server's input is the AI's translation result, and its output is the converted data sent to the terminal.

[0483] Step 7:

[0484] The terminal translates the necessary parts of the user interface based on the received translated text. This allows the user to view the information in their preferred language. The input is the translation result from the server, and the output is the translated user interface.

[0485] (Application Example 1)

[0486] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0487] In a multilingual environment, a key challenge is to mitigate the difficulties users face due to language barriers when purchasing or paying for goods, enabling them to conduct business transactions with confidence. In particular, when visiting a country or region where different languages ​​are spoken, it is essential to quickly and accurately translate information into the user's native language and provide smooth service.

[0488] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0489] In this invention, the server includes a device that stores the user's selected language settings on the terminal, a device that acquires information presented using an identification image on the terminal and generates a translation request, and a device that performs the translation process using the generated artificial intelligence. This enables users to quickly verify information in their native language and conduct transactions reliably, even in different language environments.

[0490] A "terminal" is an electronic device that a user directly operates to display information.

[0491] A "user" is defined as the entity that uses the system to translate or verify information.

[0492] "Language settings" means that the user specifies their native language or preferred language.

[0493] A "device" is a component of hardware or software designed to perform a specific function.

[0494] "Information sources" refer to websites and applications accessible through a device.

[0495] "Detection" is the act of analyzing information in order to identify a specified language.

[0496] A "translation request" is the process of requesting that information be converted from the original language to the user's desired language.

[0497] A "communication device" is a network-connected server used to send and receive data.

[0498] "Artificial intelligence" is a technology in which computers mimic human intellectual tasks and perform functions such as natural language processing.

[0499] An "identification image" is a visual element provided in the form of a code or tag for digitizing information and extracting its content.

[0500] The terminal is a device that stores the user's selected language settings and translates information based on them. When the terminal accesses an information source such as a web page or application, it first detects the display language and acquires that information as an identification image. Based on the acquired information, it generates a translation request and sends it to the communication device. This communication device functions as a server and, based on the received translation request, uses a generating AI to translate the original data. Specifically, it performs the translation using a generated artificial intelligence language model—for example, a modern neural network model. The translated content is sent back to the terminal, which displays it in the corresponding location within the user interface.

[0501] This system allows users to quickly grasp information and take appropriate action without experiencing stress, even in different language environments. For example, one scenario could involve scanning a QR code for a product sold in Japan, translating the product details into English, and then proceeding with the purchase. In this case, the prompt would be something like, "Please translate into English."

[0502] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0503] Step 1:

[0504] The device saves the language settings specified by the user. During this process, it retrieves the desired language data entered by the user through the device's language settings interface and writes it to the device's memory. This saved data is then used in subsequent processing.

[0505] Step 2:

[0506] The terminal detects the display language of the information source being accessed. It analyzes text information obtained from web pages and applications and applies a language recognition algorithm to determine the language of the input data. This process is generally performed using a natural language processing library, and the detected language code is generated as output.

[0507] Step 3:

[0508] The terminal creates a translation request based on the detected language and the language specified by the user and sends it to the communication device. In this process, it receives the detected language, target language, and text information to be translated as input, and constructs a translation request data package that includes a prompt. As output, data in a format suitable for the communication protocol is generated.

[0509] Step 4:

[0510] The server processes translation requests received from terminals and performs data translation using a generative AI model. Following the translation instructions in the request, the original text is input into the language model, and the AI ​​outputs the text result converted into the target language specified by the AI. This process utilizes advanced neural language models.

[0511] Step 5:

[0512] The server sends the translation results to the terminal. The generated translated text is formatted into a data packet and sent back to the terminal via the communication line. The output data is transmitted in an appropriate format considering the communication conditions.

[0513] Step 6:

[0514] The device reflects the received translation data in the interface and displays it to the user. The acquired translated text is applied to UI elements and output in a visually verifiable format. The user can then decide on their actions based on this information.

[0515] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0516] This invention is a multilingual system combining a terminal, a server, a generative AI, and an emotion engine. This system enables users to communicate smoothly across language barriers. The user begins using the terminal by launching an application or website and selecting their desired language. The terminal saves this language setting and automatically detects the language of the displayed content.

[0517] The device then generates a translation request and sends it to the server. This request includes the original display language, the user's selected language, and the specific text information to be translated. The server processes the received request and forwards it to the generation AI. This AI uses a highly accurate neural model to translate the text into the user's desired language.

[0518] Furthermore, the addition of an emotion engine allows the system to analyze user voice input and facial expression data to recognize the user's emotions. The server then uses the results of emotion recognition to apply the most appropriate expression to the situation through a generation AI, adjusting the translation result accordingly.

[0519] The translated and emotion-adjusted text is returned from the server to the device and displayed in the user interface in the most optimal format. As a concrete example, consider a scenario where a user uses a Japanese tourist information app and smiles with satisfaction after reading the information. The emotion engine detects positive emotions from the user's facial expressions and provides a soft, friendly translation that matches their intent. This allows the user to use the app without stress and have a better experience.

[0520] This system not only translates language but also presents content that reflects the user's emotional state, making it possible to provide information more appropriately and effectively to people from diverse cultural backgrounds.

[0521] The following describes the processing flow.

[0522] Step 1:

[0523] The user launches the application on their device and selects their desired language on the language settings screen. The device saves this language setting to the user profile.

[0524] Step 2:

[0525] The device activates a translation agent and automatically detects the display language of the website or application the user is trying to view. This detection is performed by analyzing the text and metadata on the page.

[0526] Step 3:

[0527] The device extracts the necessary text based on the detected display language and the user's selected language, and generates a translation request for the server. This request includes the specific text information that needs to be translated.

[0528] Step 4:

[0529] The server receives the translation request and sends a request to the generative AI to perform the translation. The generative AI uses a multilingual translation model to perform an accurate translation from the original language to the user's selected language.

[0530] Step 5:

[0531] The device also acquires data on the user's facial expressions and voice and sends it to the emotion engine. The emotion engine analyzes this data to identify the user's emotional state.

[0532] Step 6:

[0533] The server optimizes the translation results created by the generative AI based on emotional data from the emotion engine. It applies translation expressions that match the emotion and adjusts them so that the message is conveyed to the user in the most appropriate way.

[0534] Step 7:

[0535] The server sends the adjusted translation results to the terminal. The terminal then reflects the translated text in the user interface of the application or website and displays it to the user.

[0536] Step 8:

[0537] Through translated content, users can receive flexible information tailored to their chosen language and sentiment, enabling them to perform meaningful actions.

[0538] (Example 2)

[0539] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0540] In addition to facilitating communication between multiple languages, it is necessary to address the challenge of providing translations that take into account the user's emotional state. Conventional systems are limited to language conversion and have difficulty providing flexible translations that respond to the user's emotions and circumstances.

[0541] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0542] In this invention, the server includes means for converting character information with high accuracy based on received conversion requests, means for generating translation results using generative artificial intelligence, and means for recognizing the user's emotions using an emotion analysis device and adjusting the translation results accordingly. This makes it possible to provide flexible and natural translation results that reflect the user's emotions.

[0543] A "terminal" refers to an information processing device, a device that the user directly operates.

[0544] An "information processing device" is a part of a computer system and is a device used to process digital content.

[0545] "Digital content" is a general term for information and media that are expressed and stored electronically.

[0546] "Display language" refers to the language used by information processing devices and digital content.

[0547] A "conversion request" is a command generated and sent to a communication device for the purpose of converting languages.

[0548] A "communication device" is a part of the hardware or software used to send and receive data and information between terminals and other devices.

[0549] "Textual information" refers to linguistic data expressed in text format.

[0550] "Generative artificial intelligence" refers to artificial systems that use technologies such as machine learning and neural networks to mimic human intellectual activity.

[0551] An "emotion analysis device" is a device that analyzes a user's emotional state and provides appropriate feedback.

[0552] A "translation result" is a text in a new language format generated based on a conversion request.

[0553] A "user interface" refers to the screens and operating methods that a user uses to interact with a device or application.

[0554] This invention is a system that combines a terminal operated directly by the user, an information processing device for processing digital content, and a communication device for sending and receiving data. The user can start the system by using a designated application or website on the terminal and selecting their desired language. The terminal stores this language setting internally and uses that information in subsequent processes.

[0555] The terminal utilizes natural language processing technology to detect the language of the digital content being displayed. This identifies the basic linguistic characteristics of the content and generates a translation request. The generated translation request includes the detected language, the user-selected language, and the character information that needs to be translated. The translation request is sent to the server via a communication device.

[0556] The server uses a highly trained generative artificial intelligence model to translate incoming text information with high accuracy. During this process, an emotion analyzer analyzes the user's voice and facial expressions to identify their emotional state. Based on this emotion data, the server adjusts the translation results, going beyond simple language conversion to provide natural and appropriate expressions that reflect the user's emotions.

[0557] As a concrete example, consider a scenario where a user is using a Japanese tourist information app and wants to understand the guide information in English. If the user is smiling and looking at the screen, the emotion analyzer will detect his positive emotions, and the server can return a friendly translation accordingly.

[0558] An example of a prompt for a generative AI model would be text like, "Translate this Japanese tourist guide into English. The user is smiling." Based on this prompt, the system provides the best possible translation in the specified language.

[0559] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0560] Step 1:

[0561] The user launches an application or website on their device and selects their desired language. This input prompts the device to save the user's language settings to its data storage. This saved information is then used for subsequent language processing.

[0562] Step 2:

[0563] The terminal detects the language of the digital content being displayed. The input is content data, and natural language processing techniques are applied to analyze its linguistic characteristics. The output is information about the detected display language, which forms the basis for creating translation requests.

[0564] Step 3:

[0565] The terminal generates a translation request based on the detected display language and the language selected by the user. The input is the language detection result and the user's selected language, and the output is request data, which includes the character information to be translated. This request is sent to the server via a communication device.

[0566] Step 4:

[0567] The server uses artificial intelligence to translate text information based on the received translation request. The input is text information received from the terminal, and the AI's neural network is used to perform language conversion. The output is translated text data.

[0568] Step 5:

[0569] When a user provides voice input or displays facial expressions, the terminal processes this data using an emotion analysis device. The input consists of the user's real-time voice and visual data, and an emotion recognition algorithm is applied to identify the user's emotional state. The output is the analyzed emotion information.

[0570] Step 6:

[0571] The server adjusts the translation results based on the sentiment analysis. The input consists of the AI-generated translation results and sentiment data, and the tone and expression are adjusted to achieve natural dialogue. The output is the optimally translated text that corresponds to the sentiment.

[0572] Step 7:

[0573] The server returns the final translated text to the terminal. The output text is transmitted to the terminal via a communication device and displayed on the user interface. This allows the user to visually confirm the translation results in an appropriate format.

[0574] (Application Example 2)

[0575] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0576] In multilingual communication, it is necessary not only to convert text into text, but also to accurately convey the speaker's emotions and intentions. However, conventional systems have limitations in language translation accuracy and fail to provide appropriate translations that take into account the user's emotions. In particular, in fields such as tourist information, there is a need to provide accurate and user-friendly information that aligns with the user's feelings.

[0577] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0578] In this invention, the server includes means for converting text, means for recognizing and adjusting the user's emotions from their facial expressions and voice, and means for transmitting the converted information to a terminal. This makes it possible to provide translation results that are adapted to the user's emotions.

[0579] A "terminal" is an electronic device used by a user, and is a device for inputting and displaying information.

[0580] "Language settings" refer to settings that allow users to specify their preferred language.

[0581] An "information page" is a general term for informational content displayed on online web pages and applications.

[0582] A "symbol" is an element that represents information, such as letters or words, in natural language.

[0583] A "translation request" is request information generated to convert text into another language.

[0584] An "information processing device" is a system that processes received data and sends a response to a terminal.

[0585] "Textual information" refers to data in text format, which is information written in human language.

[0586] An "emotion analysis device" is a system for analyzing a user's emotions, and it recognizes emotions using voice and facial expression data.

[0587] A "dialogue screen" is the screen portion that displays information via the user interface.

[0588] To implement this application, a system is built using a server as the information processing device and the user's terminal as the client. The user uses a dedicated application installed on the terminal to set the desired language and then begins accessing the information page. The terminal detects the displayed symbols and sends a translation request to the server based on them.

[0589] To process incoming translation requests, the server first analyzes the text information and performs the appropriate conversion. The software used includes extracting text from images using the Google Cloud Vision API and translating the language using the Azure Translator API. Furthermore, an emotion analyzer recognizes emotions from the user's facial expressions and voice. In this process, the Face API and Emotion API from Microsoft Azure Cognitive Services are used to evaluate the user's emotional state.

[0590] Based on recognized emotions, a generative AI model is used to optimize the translation results. This generates flexible translations that are sensitive to the user's emotions. The optimized translation, using the generated prompt text, is sent to the device and displayed through the user interface.

[0591] As a concrete example, when a user takes a picture of a sign at a tourist attraction with their camera, the information processing device analyzes the image and translates it into the user's preferred language. At the same time, if it determines that the user is smiling and showing interest, it generates a translation that gently conveys detailed background information.

[0592] An example of a prompt for a generative AI model would be: "Translate the following Japanese text into English and add a detailed explanation if the user is interested: Japanese text."

[0593] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0594] Step 1:

[0595] The user launches the application on their device and sets their desired language. The input is the user's language selection, and the output is the set language information saved on the device. This language setting is used in the subsequent translation process.

[0596] Step 2:

[0597] The device accesses an information page and detects the displayed symbols within it. The input is the information page accessed by the user, and the output is the detected symbol data. This is done by extracting text from the image using the Google Cloud Vision API.

[0598] Step 3:

[0599] The terminal generates a translation request based on the detected symbol data and stored language settings, and sends it to the server. The input is the symbol data and language settings, and the output is the translation request data sent to the server. This prepares the server to begin the translation process.

[0600] Step 4:

[0601] The server analyzes the character information based on the received translation request and uses the Azure Translator API to translate the symbolic data into the user's desired language. The input is the translation request data, and the output is the translated text data. Here, data transformation for translation takes place.

[0602] Step 5:

[0603] The server uses an emotion analysis device to analyze the user's facial expressions and voice in order to understand the user's emotional state. The input is facial and voice data sent from the terminal, and the output is a recognition of the user's emotional state. Emotion analysis is performed using Microsoft Azure Cognitive Services' Face API and Emotion API.

[0604] Step 6:

[0605] The server optimizes the translation results using a generative AI model based on the recognized emotional state. The input is the emotional state and the initial translated text; the output is a flexible translation that reflects the emotion. The generative AI model makes appropriate adjustments using prompt sentences.

[0606] Step 7:

[0607] The server sends optimized translation results to the terminal, which displays the translation results through the user interface. The input is the optimized translation data received from the server, and the output is the translation result displayed to the user. This allows the user to receive translation results that are adapted to their emotions.

[0608] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0609] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0610] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0611] [Fourth Embodiment]

[0612] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0613] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0614] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0615] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0616] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0617] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0618] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0619] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0620] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0621] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0622] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0623] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0624] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0625] This invention is a multilingual system utilizing a terminal, a server, and a generative AI. The terminal used by the user simultaneously launches its own translation agent when an application or website is started. This translation agent records the language selected by the user on the terminal as initial setting data and automatically detects the display language of the content the terminal is accessing.

[0626] Based on the detected display language, the terminal generates and sends a translation request to the server. This translation request includes the original language, the target language (the language selected by the user), and the text information that needs to be translated.

[0627] The server processes the received request and forwards it to the generative AI. The generative AI uses advanced language models, such as neural network models, to translate the received text into the user's selected language. The translated text information is then sent to the terminal by the server as the translation result.

[0628] The device translates and displays the relevant parts of the application or website's user interface based on the translation results received from the server. This allows the user to view the information in their chosen language.

[0629] As a concrete example, consider a scenario where a user who only speaks English uses a Japanese public transportation timetable app. The user launches the app and sets English as the preferred language. The device's translation agent detects the timetable information provided in Japanese and sends a request to the server to translate it into English. The server uses generative AI to translate the information, and the translated timetable is displayed on the device. Based on this, the user can plan their travel schedule and enjoy a comfortable sightseeing experience.

[0630] The following describes the processing flow.

[0631] Step 1:

[0632] The user launches the application on their device and sets their preferred language. This setting information is saved on the device and used in subsequent processes.

[0633] Step 2:

[0634] The device activates a translation agent and detects the display language of the content the user is viewing. This detection is performed by analyzing the text on the page.

[0635] Step 3:

[0636] The device creates a translation request based on the detected display language and the language selected by the user, and sends it to the server. This request also includes the specific text that needs to be translated.

[0637] Step 4:

[0638] The server processes the received translation request and passes it to the generating AI. The AI ​​model translates the requested text into the user's selected language.

[0639] Step 5:

[0640] The server sends the translated result to the terminal. This result is data containing the translated text.

[0641] Step 6:

[0642] The device translates the user interface of an application or website based on the received translation results and displays it to the user.

[0643] Step 7:

[0644] Users can perform the necessary actions using translated content, thereby enabling them to accurately understand and utilize the information.

[0645] (Example 1)

[0646] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0647] When accessing information in different languages, users often face challenges in understanding the information, particularly in the interface, which can make smooth operation difficult. This means that language barriers are a significant obstacle to information utilization and international communication.

[0648] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0649] In this invention, the server includes means for converting character data based on a received translation request, means for transmitting the conversion result to a terminal, and means for receiving translation requests from the terminal. This enables smooth translation and display of information in different languages.

[0650] A "terminal" is an electronic device that users directly operate to access information.

[0651] "Language settings" is a component that allows users to specify their preferred language for displaying and interacting with information.

[0652] "Display language" refers to the language used when information or an interface is presented visually to the user.

[0653] "Means of recognition" refers to a function that allows the terminal to automatically identify the language in which information is displayed.

[0654] A "translation request" is a request sent from a terminal to a server that includes the source language and the target language of the information.

[0655] A "device for aggregating information" is a central computer that performs data conversion and management based on requests from terminals.

[0656] "Means of converting text data" refers to the process of automatically translating information into another language.

[0657] "Conversion result" refers to the output after converting character data to another language.

[0658] "User interface" refers to the screen that the user directly interacts with, where information is presented through the terminal.

[0659] "Generative artificial intelligence" is a technology that uses neural network models to automatically translate text into other languages.

[0660] This invention is a system that facilitates multilingual information access and includes a terminal for user access, a server for aggregating information, and a generative AI with advanced translation capabilities.

[0661] The user launches a specific source of information (e.g., a webpage or application) on their device. This device has a feature that allows the user to specify their preferred display language as the default setting. The device automatically parses HTML tags and metadata to recognize the language of the displayed information.

[0662] The device generates a translation request to a server that aggregates information, based on its recognized language settings. This request includes the original language, the user's desired target language, and the text to be translated. The server receives this request and sends it as a prompt to the artificial intelligence that generates the information.

[0663] The generative AI uses a neural network model to translate text into the desired language. This AI considers context and cultural nuances to produce natural-sounding translations. The translated text is then sent to the terminal via a server.

[0664] The terminal displays the received conversion result in the user's preferred language on the interface they interact with. This allows the user to understand and interact with information in their chosen language.

[0665] As a concrete example, consider a scenario where a user who only understands English accesses a tourist information webpage provided in Japanese. In this case, the user opens the webpage and specifies English as the desired display language. The device recognizes the Japanese content and sends a request to the server to translate it into English. The server uses generative AI to translate the information, and the English content is displayed on the device. This allows the user to use the information without being aware of the language barrier.

[0666] An example of a prompt message would be the text format: "Please translate the following Japanese sentence into English: 'Hello'".

[0667] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0668] Step 1:

[0669] The user launches a specific application or website using the device. The translation agent on the device records the language the user initially selected. Based on this input, the device retains the user's preferred language as output.

[0670] Step 2:

[0671] The device automatically detects the display language of the content being accessed. This process involves analyzing HTML tags and metadata, for example, by checking the "lang" attribute to identify the language. The detected language is then used as input data for the translation process.

[0672] Step 3:

[0673] The device generates a translation request based on the detected display language and the user's language setting, and sends it to the server. This translation request includes the source language, the target language, and the text to be translated. This is then sent to the server as output data for the translation process.

[0674] Step 4:

[0675] The server analyzes the translation request received from the terminal and forwards the request to the generating AI model along with the generated prompt. The server's input is the translation request, and its output is the AI ​​request, which includes the prompt.

[0676] Step 5:

[0677] The generative AI processes data based on the received prompt to produce a natural translation. The AI ​​uses a neural network to analyze context and nuances, and outputs the translated text. The output is the translated text in the language selected by the user.

[0678] Step 6:

[0679] The server receives the translated text from the generating AI and sends it to the user's terminal. The server's input is the AI's translation result, and its output is the converted data sent to the terminal.

[0680] Step 7:

[0681] The terminal translates the necessary parts of the user interface based on the received translated text. This allows the user to view the information in their preferred language. The input is the translation result from the server, and the output is the translated user interface.

[0682] (Application Example 1)

[0683] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0684] In a multilingual environment, a key challenge is to mitigate the difficulties users face due to language barriers when purchasing or paying for goods, enabling them to conduct business transactions with confidence. In particular, when visiting a country or region where different languages ​​are spoken, it is essential to quickly and accurately translate information into the user's native language and provide smooth service.

[0685] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0686] In this invention, the server includes a device that stores the user's selected language settings on the terminal, a device that acquires information presented using an identification image on the terminal and generates a translation request, and a device that performs the translation process using the generated artificial intelligence. This enables users to quickly verify information in their native language and conduct transactions reliably, even in different language environments.

[0687] A "terminal" is an electronic device that a user directly operates to display information.

[0688] A "user" is defined as the entity that uses the system to translate or verify information.

[0689] "Language settings" means that the user specifies their native language or preferred language.

[0690] A "device" is a component of hardware or software designed to perform a specific function.

[0691] "Information sources" refer to websites and applications accessible through a device.

[0692] "Detection" is the act of analyzing information in order to identify a specified language.

[0693] A "translation request" is the process of requesting that information be converted from the original language to the user's desired language.

[0694] A "communication device" is a network-connected server used to send and receive data.

[0695] "Artificial intelligence" is a technology in which computers mimic human intellectual tasks and perform functions such as natural language processing.

[0696] An "identification image" is a visual element provided in the form of a code or tag for digitizing information and extracting its content.

[0697] The terminal is a device that stores the user's selected language settings and translates information based on them. When the terminal accesses an information source such as a web page or application, it first detects the display language and acquires that information as an identification image. Based on the acquired information, it generates a translation request and sends it to the communication device. This communication device functions as a server and, based on the received translation request, uses a generating AI to translate the original data. Specifically, it performs the translation using a generated artificial intelligence language model—for example, a modern neural network model. The translated content is sent back to the terminal, which displays it in the corresponding location within the user interface.

[0698] This system allows users to quickly grasp information and take appropriate action without experiencing stress, even in different language environments. For example, one scenario could involve scanning a QR code for a product sold in Japan, translating the product details into English, and then proceeding with the purchase. In this case, the prompt would be something like, "Please translate into English."

[0699] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0700] Step 1:

[0701] The device saves the language settings specified by the user. During this process, it retrieves the desired language data entered by the user through the device's language settings interface and writes it to the device's memory. This saved data is then used in subsequent processing.

[0702] Step 2:

[0703] The terminal detects the display language of the information source being accessed. It analyzes text information obtained from web pages and applications and applies a language recognition algorithm to determine the language of the input data. This process is generally performed using a natural language processing library, and the detected language code is generated as output.

[0704] Step 3:

[0705] The terminal creates a translation request based on the detected language and the language specified by the user and sends it to the communication device. In this process, it receives the detected language, target language, and text information to be translated as input, and constructs a translation request data package that includes a prompt. As output, data in a format suitable for the communication protocol is generated.

[0706] Step 4:

[0707] The server processes translation requests received from terminals and performs data translation using a generative AI model. Following the translation instructions in the request, the original text is input into the language model, and the AI ​​outputs the text result converted into the target language specified by the AI. This process utilizes advanced neural language models.

[0708] Step 5:

[0709] The server sends the translation results to the terminal. The generated translated text is formatted into a data packet and sent back to the terminal via the communication line. The output data is transmitted in an appropriate format considering the communication conditions.

[0710] Step 6:

[0711] The device reflects the received translation data in the interface and displays it to the user. The acquired translated text is applied to UI elements and output in a visually verifiable format. The user can then decide on their actions based on this information.

[0712] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0713] This invention is a multilingual system combining a terminal, a server, a generative AI, and an emotion engine. This system enables users to communicate smoothly across language barriers. The user begins using the terminal by launching an application or website and selecting their desired language. The terminal saves this language setting and automatically detects the language of the displayed content.

[0714] The device then generates a translation request and sends it to the server. This request includes the original display language, the user's selected language, and the specific text information to be translated. The server processes the received request and forwards it to the generation AI. This AI uses a highly accurate neural model to translate the text into the user's desired language.

[0715] Furthermore, the addition of an emotion engine allows the system to analyze user voice input and facial expression data to recognize the user's emotions. The server then uses the results of emotion recognition to apply the most appropriate expression to the situation through a generation AI, adjusting the translation result accordingly.

[0716] The translated and emotion-adjusted text is returned from the server to the device and displayed in the user interface in the most optimal format. As a concrete example, consider a scenario where a user uses a Japanese tourist information app and smiles with satisfaction after reading the information. The emotion engine detects positive emotions from the user's facial expressions and provides a soft, friendly translation that matches their intent. This allows the user to use the app without stress and have a better experience.

[0717] This system not only translates language but also presents content that reflects the user's emotional state, making it possible to provide information more appropriately and effectively to people from diverse cultural backgrounds.

[0718] The following describes the processing flow.

[0719] Step 1:

[0720] The user launches the application on their device and selects their desired language on the language settings screen. The device saves this language setting to the user profile.

[0721] Step 2:

[0722] The device activates a translation agent and automatically detects the display language of the website or application the user is trying to view. This detection is performed by analyzing the text and metadata on the page.

[0723] Step 3:

[0724] The device extracts the necessary text based on the detected display language and the user's selected language, and generates a translation request for the server. This request includes the specific text information that needs to be translated.

[0725] Step 4:

[0726] The server receives the translation request and sends a request to the generative AI to perform the translation. The generative AI uses a multilingual translation model to perform an accurate translation from the original language to the user's selected language.

[0727] Step 5:

[0728] The device also acquires data on the user's facial expressions and voice and sends it to the emotion engine. The emotion engine analyzes this data to identify the user's emotional state.

[0729] Step 6:

[0730] The server optimizes the translation results created by the generative AI based on emotional data from the emotion engine. It applies translation expressions that match the emotion and adjusts them so that the message is conveyed to the user in the most appropriate way.

[0731] Step 7:

[0732] The server sends the adjusted translation results to the terminal. The terminal then reflects the translated text in the user interface of the application or website and displays it to the user.

[0733] Step 8:

[0734] Through translated content, users can receive flexible information tailored to their chosen language and sentiment, enabling them to perform meaningful actions.

[0735] (Example 2)

[0736] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0737] In addition to facilitating communication between multiple languages, it is necessary to address the challenge of providing translations that take into account the user's emotional state. Conventional systems are limited to language conversion and have difficulty providing flexible translations that respond to the user's emotions and circumstances.

[0738] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0739] In this invention, the server includes means for converting character information with high accuracy based on received conversion requests, means for generating translation results using generative artificial intelligence, and means for recognizing the user's emotions using an emotion analysis device and adjusting the translation results accordingly. This makes it possible to provide flexible and natural translation results that reflect the user's emotions.

[0740] A "terminal" refers to an information processing device, a device that the user directly operates.

[0741] An "information processing device" is a part of a computer system and is a device used to process digital content.

[0742] "Digital content" is a general term for information and media that are expressed and stored electronically.

[0743] "Display language" refers to the language used by information processing devices and digital content.

[0744] A "conversion request" is a command generated and sent to a communication device for the purpose of converting languages.

[0745] A "communication device" is a part of the hardware or software used to send and receive data and information between terminals and other devices.

[0746] "Textual information" refers to linguistic data expressed in text format.

[0747] "Generative artificial intelligence" refers to artificial systems that use technologies such as machine learning and neural networks to mimic human intellectual activity.

[0748] An "emotion analysis device" is a device that analyzes a user's emotional state and provides appropriate feedback.

[0749] A "translation result" is a text in a new language format generated based on a conversion request.

[0750] A "user interface" refers to the screens and operating methods that a user uses to interact with a device or application.

[0751] This invention is a system that combines a terminal operated directly by the user, an information processing device for processing digital content, and a communication device for sending and receiving data. The user can start the system by using a designated application or website on the terminal and selecting their desired language. The terminal stores this language setting internally and uses that information in subsequent processes.

[0752] The terminal utilizes natural language processing technology to detect the language of the digital content being displayed. This identifies the basic linguistic characteristics of the content and generates a translation request. The generated translation request includes the detected language, the user-selected language, and the character information that needs to be translated. The translation request is sent to the server via a communication device.

[0753] The server uses a highly trained generative artificial intelligence model to translate incoming text information with high accuracy. During this process, an emotion analyzer analyzes the user's voice and facial expressions to identify their emotional state. Based on this emotion data, the server adjusts the translation results, going beyond simple language conversion to provide natural and appropriate expressions that reflect the user's emotions.

[0754] As a concrete example, consider a scenario where a user is using a Japanese tourist information app and wants to understand the guide information in English. If the user is smiling and looking at the screen, the emotion analyzer will detect his positive emotions, and the server can return a friendly translation accordingly.

[0755] An example of a prompt for a generative AI model would be text like, "Translate this Japanese tourist guide into English. The user is smiling." Based on this prompt, the system provides the best possible translation in the specified language.

[0756] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0757] Step 1:

[0758] The user launches an application or website on their device and selects their desired language. This input prompts the device to save the user's language settings to its data storage. This saved information is then used for subsequent language processing.

[0759] Step 2:

[0760] The terminal detects the language of the digital content being displayed. The input is content data, and natural language processing techniques are applied to analyze its linguistic characteristics. The output is information about the detected display language, which forms the basis for creating translation requests.

[0761] Step 3:

[0762] The terminal generates a translation request based on the detected display language and the language selected by the user. The input is the language detection result and the user's selected language, and the output is request data, which includes the character information to be translated. This request is sent to the server via a communication device.

[0763] Step 4:

[0764] The server uses artificial intelligence to translate text information based on the received translation request. The input is text information received from the terminal, and the AI's neural network is used to perform language conversion. The output is translated text data.

[0765] Step 5:

[0766] When a user provides voice input or displays facial expressions, the terminal processes this data using an emotion analysis device. The input consists of the user's real-time voice and visual data, and an emotion recognition algorithm is applied to identify the user's emotional state. The output is the analyzed emotion information.

[0767] Step 6:

[0768] The server adjusts the translation results based on the sentiment analysis. The input consists of the AI-generated translation results and sentiment data, and the tone and expression are adjusted to achieve natural dialogue. The output is the optimally translated text that corresponds to the sentiment.

[0769] Step 7:

[0770] The server returns the final translated text to the terminal. The output text is transmitted to the terminal via a communication device and displayed on the user interface. This allows the user to visually confirm the translation results in an appropriate format.

[0771] (Application Example 2)

[0772] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0773] In multilingual communication, it is necessary not only to convert text into text, but also to accurately convey the speaker's emotions and intentions. However, conventional systems have limitations in language translation accuracy and fail to provide appropriate translations that take into account the user's emotions. In particular, in fields such as tourist information, there is a need to provide accurate and user-friendly information that aligns with the user's feelings.

[0774] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0775] In this invention, the server includes means for converting text, means for recognizing and adjusting the user's emotions from their facial expressions and voice, and means for transmitting the converted information to a terminal. This makes it possible to provide translation results that are adapted to the user's emotions.

[0776] A "terminal" is an electronic device used by a user, and is a device for inputting and displaying information.

[0777] "Language settings" refer to settings that allow users to specify their preferred language.

[0778] An "information page" is a general term for informational content displayed on online web pages and applications.

[0779] A "symbol" is an element that represents information, such as letters or words, in natural language.

[0780] A "translation request" is request information generated to convert text into another language.

[0781] An "information processing device" is a system that processes received data and sends a response to a terminal.

[0782] "Textual information" refers to data in text format, which is information written in human language.

[0783] An "emotion analysis device" is a system for analyzing a user's emotions, and it recognizes emotions using voice and facial expression data.

[0784] A "dialogue screen" is the screen portion that displays information via the user interface.

[0785] To implement this application, a system is built using a server as the information processing device and the user's terminal as the client. The user uses a dedicated application installed on the terminal to set the desired language and then begins accessing the information page. The terminal detects the displayed symbols and sends a translation request to the server based on them.

[0786] To process incoming translation requests, the server first analyzes the text information and performs the appropriate conversion. The software used includes extracting text from images using the Google Cloud Vision API and translating the language using the Azure Translator API. Furthermore, an emotion analyzer recognizes emotions from the user's facial expressions and voice. In this process, the Face API and Emotion API from Microsoft Azure Cognitive Services are used to evaluate the user's emotional state.

[0787] Based on recognized emotions, a generative AI model is used to optimize the translation results. This generates flexible translations that are sensitive to the user's emotions. The optimized translation, using the generated prompt text, is sent to the device and displayed through the user interface.

[0788] As a concrete example, when a user takes a picture of a sign at a tourist attraction with their camera, the information processing device analyzes the image and translates it into the user's preferred language. At the same time, if it determines that the user is smiling and showing interest, it generates a translation that gently conveys detailed background information.

[0789] An example of a prompt for a generative AI model would be: "Translate the following Japanese text into English and add a detailed explanation if the user is interested: Japanese text."

[0790] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0791] Step 1:

[0792] The user launches the application on their device and sets their desired language. The input is the user's language selection, and the output is the set language information saved on the device. This language setting is used in the subsequent translation process.

[0793] Step 2:

[0794] The device accesses an information page and detects the displayed symbols within it. The input is the information page accessed by the user, and the output is the detected symbol data. This is done by extracting text from the image using the Google Cloud Vision API.

[0795] Step 3:

[0796] The terminal generates a translation request based on the detected symbol data and stored language settings, and sends it to the server. The input is the symbol data and language settings, and the output is the translation request data sent to the server. This prepares the server to begin the translation process.

[0797] Step 4:

[0798] The server analyzes the character information based on the received translation request and uses the Azure Translator API to translate the symbolic data into the user's desired language. The input is the translation request data, and the output is the translated text data. Here, data transformation for translation takes place.

[0799] Step 5:

[0800] The server uses an emotion analysis device to analyze the user's facial expressions and voice in order to understand the user's emotional state. The input is facial and voice data sent from the terminal, and the output is a recognition of the user's emotional state. Emotion analysis is performed using Microsoft Azure Cognitive Services' Face API and Emotion API.

[0801] Step 6:

[0802] The server optimizes the translation results using a generative AI model based on the recognized emotional state. The input is the emotional state and the initial translated text; the output is a flexible translation that reflects the emotion. The generative AI model makes appropriate adjustments using prompt sentences.

[0803] Step 7:

[0804] The server sends optimized translation results to the terminal, which displays the translation results through the user interface. The input is the optimized translation data received from the server, and the output is the translation result displayed to the user. This allows the user to receive translation results that are adapted to their emotions.

[0805] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0806] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0807] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0808] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0809] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0810] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0811] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0812] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0813] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0814] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0815] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0816] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0817] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0818] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0819] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0820] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0821] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0822] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0823] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0824] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0825] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0826] The following is further disclosed regarding the embodiments described above.

[0827] (Claim 1)

[0828] A means by which the device saves the user's selected language settings,

[0829] A means for detecting the display language of the website or application that the device is accessing,

[0830] A means for generating a translation request based on the language detected by the terminal and the language selected by the user, and sending it to the server,

[0831] A means for translating text based on a translation request received by the server,

[0832] A means by which the server sends the translation result to the terminal,

[0833] A means of displaying the translation results received by the terminal,

[0834] A multilingual system including

[0835] (Claim 2)

[0836] The system according to claim 1, wherein the translation result is applied to a specific part of the user interface.

[0837] (Claim 3)

[0838] The system according to claim 1, wherein the translation process is performed using a generative AI.

[0839] "Example 1"

[0840] (Claim 1)

[0841] A means for the device to record the language settings selected by the user,

[0842] A means for the terminal to recognize the display language of the information it is accessing,

[0843] A means for creating a translation request based on the language recognized by the terminal and the language selected by the user, and sending it to a device that aggregates the information,

[0844] A device that aggregates information has means for converting text data based on received translation requests,

[0845] A means by which a device that aggregates information transmits the conversion result to a terminal,

[0846] A means of displaying the conversion result received by the terminal,

[0847] A system that includes this.

[0848] (Claim 2)

[0849] The system according to claim 1, wherein the conversion result is applied to a specific portion of the interface that the user comes into contact with.

[0850] (Claim 3)

[0851] The system according to claim 1, which is performed using artificial intelligence generated by the conversion process.

[0852] "Application Example 1"

[0853] (Claim 1)

[0854] A device that saves the user's selected language settings,

[0855] A device that detects the display language of the information source being accessed by the terminal,

[0856] A device that generates a translation request based on the detected language and the language selected by the user, and transmits it to a communication device,

[0857] A device that translates data based on a translation request received by a communication device,

[0858] A communication device that transmits the translation result to the terminal,

[0859] A device that displays the received translation results on the terminal,

[0860] A device that uses an identification image to obtain information presented by a terminal and generates a translation request,

[0861] A system that includes this.

[0862] (Claim 2)

[0863] The system according to claim 1, wherein the translation result is applied to a specific part of the user interface.

[0864] (Claim 3)

[0865] The system according to claim 1, wherein the translation process is performed using artificial intelligence that has been generated.

[0866] "Example 2 of combining an emotion engine"

[0867] (Claim 1)

[0868] A means by which the terminal saves the language settings of the information processing device,

[0869] A means by which a terminal detects the display language of digital content on an information processing device,

[0870] A means for generating a conversion request based on the language detected by the terminal and the language selected by the information processing device, and transmitting it to a communication device,

[0871] A means for converting character information based on a conversion request received by a communication device,

[0872] A means of converting text information with high accuracy using artificial intelligence generated by a communication device,

[0873] A means by which a communication device transmits the conversion result to a terminal,

[0874] A means by which a terminal recognizes the user's emotions using an emotion analysis device and adjusts the translation result based on the recognized emotions,

[0875] A means for displaying the conversion result received by the terminal,

[0876] A system that includes this.

[0877] (Claim 2)

[0878] The system according to claim 1, wherein the conversion result is applied to a specific portion of the user interface.

[0879] (Claim 3)

[0880] The system according to claim 1, comprising using an emotion analysis device to recognize the user's emotions and adjusting the translation results based on those emotions.

[0881] "Application example 2 when combining with an emotional engine"

[0882] (Claim 1)

[0883] A means by which the device saves the user's selected language settings,

[0884] A means for detecting the display symbols of the information page or application program that the terminal is accessing,

[0885] A means for generating a translation request based on the symbols detected by the terminal and the symbols selected by the user, and transmitting it to an information processing device,

[0886] A means for converting text information based on a translation request received by an information processing device,

[0887] A means by which an emotion analysis device recognizes emotions from the user's facial expressions and voice, and adjusts the translation results based on that,

[0888] A means by which the information processing device transmits the adjusted translation result to the terminal,

[0889] A means of displaying the translation results received by the terminal,

[0890] A system that includes the means to do so.

[0891] (Claim 2)

[0892] The system according to claim 1, wherein the translation result is applied to a specific portion of the user's dialogue screen.

[0893] (Claim 3)

[0894] The system according to claim 1, wherein the translation process is performed using generative AI technology. [Explanation of Symbols]

[0895] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A device that saves the user's selected language settings, A device that detects the display language of the information source being accessed by the terminal, A device that generates a translation request based on the detected language and the language selected by the user, and transmits it to a communication device, A device that translates data based on a translation request received by a communication device, A communication device that transmits the translation result to the terminal, A device that displays the received translation results on the terminal, A device that uses an identification image to obtain information presented by a terminal and generates a translation request, A system that includes this.

2. The system according to claim 1, wherein the translation result is applied to a specific part of the user interface.

3. The system according to claim 1, wherein the translation process is performed using artificial intelligence that has been generated.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A