System
The multilingual translation system addresses translation quality and speed issues by using image processing and generative AI for manga, allowing fans to access diverse content and publishers to expand efficiently.
Patent Information
- Application Number
- JP2024121521
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional manga translation methods face challenges with translation quality and speed, making it inconvenient for overseas fans and costly for domestic publishers to expand globally.
A multilingual translation system incorporating image processing, character recognition, and language translation using generative artificial intelligence to provide high-quality, context-appropriate translations.
Enables overseas fans to enjoy a wide range of manga quickly and domestically publishers to expand globally at low cost with accurate and contextually appropriate translations.
Smart Images

Figure 2026019773000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional manga translation methods have problems with translation quality and speed, making them particularly inconvenient for overseas fans and domestic publishers. Specifically, overseas fans can only read a limited number of works, and the delays and low quality of translations force them to use pirated sites. Meanwhile, for domestic publishers, global expansion is expensive and difficult to implement. For this reason, there is a need for a method to provide multilingual translations quickly and at low cost. [Means for solving the problem]
[0005] The present invention solves the aforementioned problems by providing a multilingual translation system including an image processing means, a character recognition means, a language translation means, and a translation result output means. The image processing means distinguishes between pictures and text from image files, and the character recognition means extracts text using optical character recognition technology. The language translation means performs translation using generative artificial intelligence, and the translation result output means displays the translation results on a user interface. The image processing means is also tailored to specific genres and styles, enabling high-quality translations that fit the context of the manga. This system allows overseas fans to quickly enjoy a wide range of works with high quality, and enables domestic publishers to easily expand globally at low cost.
[0006] "Image processing means" refers to a device or program that includes the function of distinguishing between pictures and text in an image file.
[0007] "Character recognition means" refers to a device or program that includes the functionality to extract characters from an image using optical character recognition technology.
[0008] "Language translation means" refers to a device or program that uses generative artificial intelligence to convert extracted characters into a target language.
[0009] The "translation result output means" refers to a device or program for displaying the translated characters to the user.
[0010] A "multilingual translation system" is a system that includes an image processing means, a character recognition means, a language translation means, and a translation result output means, and that translates into different languages.
[0011] "Optical character recognition technology" is a technology that reads characters from an image and converts them into digital text information.
[0012] "Generative artificial intelligence" refers to a machine learning model that generates or translates natural-sounding sentences based on given input data.
[0013] "User interface" refers to the screen or device through which a user interacts with a system.
[0014] A "genre" is a particular classification or category of manga works.
[0015] "Style" refers to a particular design or drawing technique used in manga works. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] The present invention provides a multilingual translation system, which includes an image processing unit, a character recognition unit, a language translation unit, and a translation result output unit.
[0038] First, the user inputs the manga image file (e.g., "manga_page.jpg") and the target language (e.g., English "en") into the terminal. The terminal receives these inputs and sends the image file path and the target language to the server.
[0039] Next, the server receives the request from the terminal, reads the image file using image processing means, and distinguishes between pictures and text. Specifically, the server converts the image file to grayscale and performs preprocessing to remove noise.
[0040] The server then uses a character recognition tool to extract characters from the processed image. Optical character recognition technology is used to convert the characters in the image into digital text data. This extracted string of characters is temporarily stored.
[0041] The server then uses a language translation tool to translate the extracted characters into the target language. Generative artificial intelligence is used to convert the extracted text into the target language. This procedure allows for a more natural and context-appropriate translation.
[0042] The translated text is finally displayed on the user interface using the translation result output means. The terminal receives a response from the server and displays the translation result to the user. This process allows the user to obtain a fast and accurate translation result.
[0043] As a concrete example, suppose a user wants to translate a Japanese manga page "manga_page.jpg" into English. When the user enters the image file path and the target language "en" into the device, the device sends this information to the server. The server processes the image, extracts the string "Hello" and translates it to "Hello". Finally, the device displays the translation result "Hello" to the user.
[0044] This allows overseas manga fans to quickly enjoy a wide range of works in high quality, while domestic publishers can easily expand globally at low cost. Because the system is tailored to specific genres and styles, it is possible to provide appropriate translations that fit the context of each work.
[0045] The processing flow will be explained below.
[0046] Step 1:
[0047] The user launches the application on their device and enters the image file path of the manga they want to translate (e.g., "manga_page.jpg") and the target language (e.g., English "en").
[0048] Step 2:
[0049] The terminal receives the user's input data and sends it to the server as a request, which includes the image file path and the target language.
[0050] Step 3:
[0051] The server receives the request from the device and analyzes the image file path and target language.
[0052] Step 4:
[0053] The server starts the image processing means and reads the specified image file (e.g., "manga_page.jpg"). It converts the image file to grayscale using OpenCV and performs preprocessing to remove noise.
[0054] Step 5:
[0055] The server distinguishes between pictures and text from the preprocessed images and identifies areas containing text.
[0056] Step 6:
[0057] The server uses a character recognition means to extract characters from the pre-processed image using optical character recognition (OCR) techniques, for example, the text "hello" is extracted.
[0058] Step 7:
[0059] The server temporarily stores the extracted character data.
[0060] Step 8:
[0061] The server uses a language translation means to translate the extracted text into the target language using generative artificial intelligence, for example, "hello" is translated into "hello".
[0062] Step 9:
[0063] The server passes the translation result to the translation result output means.
[0064] Step 10:
[0065] The translation result output means formats the translation result and converts it into a data format for display on a user interface.
[0066] Step 11:
[0067] The server generates a response including the translation result and sends it to the terminal.
[0068] Step 12:
[0069] The terminal receives the response from the server and displays the translation result (e.g., "Hello") to the user through the user interface.
[0070] Example 1
[0071] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0072] Conventional multilingual translation systems have had difficulty properly recognizing text in images and providing natural, context-appropriate translations. Furthermore, building systems that enable end users to receive fast, accurate translation results is complex and expensive.
[0073] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0074] In this invention, the server includes an image processing means, a character recognition means, and a language translation means. This makes it possible to distinguish between pictures and text from image files, extract text using optical character recognition technology, and translate the extracted text into a target language using generative artificial intelligence. Specifically, by linking a terminal where a user inputs an image file and the target language with a server that receives and processes information from the terminal, a multilingual translation system that provides fast and accurate translation results is realized.
[0075] The "image processing means" is a means for reading an image file, distinguishing between pictures and text, and performing preprocessing such as grayscale conversion and noise removal as necessary.
[0076] "Character recognition means" refers to means that uses optical character recognition technology to extract characters from an image and convert them into digital text data.
[0077] The "language translation means" is a means for translating extracted text data into a target language using generative artificial intelligence.
[0078] The "translation result output means" is a means for displaying the translated text on a user interface.
[0079] A "terminal" is a device through which a user inputs an image file and a target language and sends that information to a server.
[0080] A "server" is a computer system that performs image processing, character recognition, and language translation based on information received from a terminal, and then sends the results to the terminal.
[0081] The present invention provides a multilingual translation system, which includes an image processing unit, a character recognition unit, a language translation unit, and a translation result output unit, and further includes a terminal where a user inputs an image file and a target language, and a server that processes information received from the terminal.
[0082] First, the user inputs the manga image file (e.g., "manga_page.jpg") and the target language (e.g., English "en") into the terminal. The terminal receives these inputs and sends the image file path and the target language to the server. The image file is usually a JPEG or PNG file, and the target language is based on the user's preference.
[0083] Next, the server receives the request from the device and uses image processing to read the image file and distinguish between pictures and text. Specifically, the server uses an image processing library such as OpenCV to convert the image file to grayscale and apply a noise reduction filter such as Gaussian blur. This preprocessing improves the accuracy of character recognition.
[0084] The server then uses character recognition to extract characters from the processed image. Optical character recognition software, such as Tesseract OCR, is used to convert the characters in the image into digital text data. This extracted string is then temporarily stored.
[0085] The server then uses a language translation tool to translate the extracted characters into the target language. Generative artificial intelligence (AI) is used to convert the extracted text into the target language. This procedure allows for a more natural and contextual translation. For example, a prompt sentence such as "Please translate the following Japanese text into English: 'Hello'" is sent to the generative AI model, and the translation result is "Hello."
[0086] Finally, the translated text is displayed on the user interface using the translation result output means. The terminal receives a response from the server and displays the translation result to the user. This process allows the user to obtain a fast and accurate translation result. This system allows, for example, overseas manga fans to enjoy Japanese manga in English, and domestic publishers to easily expand globally at low cost.
[0087] As a concrete example, if a user wants to translate a manga page "manga_page.jpg" into English, the user inputs the image file path and the target language "en" into their device. The device sends this information to the server, which performs image preprocessing, character recognition, and translation. As a result, the translated text "Hello" is displayed on the user's device. By utilizing a generative AI model, this system can provide natural-sounding translations that fit the context.
[0088] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0089] Step 1:
[0090] The user inputs the manga image file and target language into the device. Specifically, the user uses the device interface to select the path of the image file (e.g., "manga_page.jpg") and specify the target language (e.g., "en"). This inputs the image file path and target language information into the device.
[0091] Input: Image file (manga_page.jpg), target language (en)
[0092] Output: Image file path and target language information
[0093] Step 2:
[0094] The device sends the input information to the server. The device packages the image file path and target language information as an HTTP request and sends it to the server. Specifically, the information is included in the HTTP request body in JSON format or multipart form data.
[0095] Input: Image file path and target language information
[0096] Output: HTTP request sent to the server
[0097] Step 3:
[0098] The server preprocesses the image file. The server reads the image file received from the device, converts the image to grayscale using OpenCV, and removes noise using Gaussian blur, etc. This improves the accuracy of character recognition.
[0099] Input: HTTP request (image file path, target language information)
[0100] Output: Preprocessed image
[0101] Step 4:
[0102] The server performs character recognition. The server passes the preprocessed image to optical character recognition software such as Tesseract OCR, which extracts characters from the image. Specifically, it converts a string of characters, such as "hello," into digital text data and temporarily stores it.
[0103] Input: Preprocessed image
[0104] Output: Extracted text data (e.g. "Hello")
[0105] Step 5:
[0106] The server translates the characters into the target language. The server uses generative artificial intelligence (AI) to convert the extracted text data into the target language. The specific prompt is "Please translate the following Japanese text into English: 'Hello'". The translation result is "Hello".
[0107] Input: Extracted text data (e.g. "Hello")
[0108] Output: Translated text data (e.g. "Hello")
[0109] Step 6:
[0110] The server sends the translation results to the device. The server packages the translated text data in JSON format and sends it to the device as an HTTP response.
[0111] Input: Translated text data (e.g. "Hello")
[0112] Output: HTTP response sent to the device
[0113] Step 7:
[0114] The terminal displays the translation result to the user. The terminal analyzes the HTTP response received from the server and displays the translation result in the user interface. Specifically, it uses a text view or an alert box to allow the user to check the translation result (e.g., "Hello").
[0115] Input: HTTP response (translated text data)
[0116] Output: The translation result displayed in the user interface (e.g. "Hello")
[0117] (Application example 1)
[0118] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0119] In today's brick-and-mortar stores, store clerks need to be proficient in multiple languages to provide multilingual support, but there are limitations to this. In particular, in areas with a large number of tourists and foreign residents, fast and accurate translation is required, but this is difficult to do manually. Therefore, there is a need for a system that allows store clerks to instantly understand product labels and information signs written in foreign languages.
[0120] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0121] In this invention, the server includes an image processing unit, a character recognition unit, and a language translation unit. This allows images captured using the smart glasses to be translated in real time and the results to be visually displayed. This allows store clerks to quickly and accurately understand foreign language information, dramatically improving multilingual support.
[0122] "Image processing means" refers to a device or device function that reads an image file and performs preprocessing such as grayscale conversion and noise removal.
[0123] "Character recognition means" refers to a function that converts characters in an image into digital text data using optical character recognition (OCR) technology.
[0124] "Language translation means" refers to a device or software that has the function of translating extracted character data into a target language.
[0125] "Translation result output means" refers to a device or a function of a device for displaying translated text data to a user.
[0126] "Smart glasses" refers to a wearable device that has a built-in camera and display and has the ability to display information in real time.
[0127] "Means for real-time translation of images captured by smart glasses" refers to a device or software that has the function of instantly processing images captured by smart glasses and displaying the translation results.
[0128] This invention is a system that performs real-time multilingual translation in a physical store using smart glasses worn by store clerks. Specific embodiments of this system are described below.
[0129] System Overview
[0130] The system includes an image processing unit, a character recognition unit, a language translation unit, and a translation result output unit, as well as smart glasses and a unit for translating images captured by the glasses in real time.
[0131] Basic operation
[0132] The server analyzes the image entered by the user and generates the translation result using image processing, character recognition, and language translation. This series of processes uses the following hardware and software:
[0133] 1. Image processing means:
[0134] Hardware: Camera built into smart glasses
[0135] Software: OpenCV (image grayscale conversion and noise reduction)
[0136] 2. Character recognition means:
[0137] Software: Tesseract OCR (Optical Character Recognition Technology)
[0138] 3. Language Translation Methods:
[0139] Software: GoogleTrans API (language translation)
[0140] 4. Translation result output method:
[0141] Hardware: Smart glasses display
[0142] Software: Custom application for displaying translation results
[0143] Examples of data processing and data calculation
[0144] When a user (store clerk) wears smart glasses and sees a product label written in a foreign language, they take a picture of the label with their camera. The captured image is processed by the smart glasses' processor, which removes noise and converts it to grayscale. Next, Tesseract OCR is used to convert the characters in the image into digital text. This text data is then translated into the desired language using the Google Translate API. The final translation result is visually displayed on the smart glasses' display, allowing the user to instantly understand the content.
[0145] Specific examples
[0146] For example, suppose a user wears smart glasses and looks at the manga page "manga_page.jpg". The camera in the smart glasses captures this image and processes it based on the target language specified by the user (e.g., English "en"). OCR technology extracts the string "Hello", which is then translated into "Hello" by the GoogleTrans API. Finally, the translation result "Hello" is displayed on the display of the smart glasses.
[0147] Prompt Sentence Examples
[0148] "Please tell us more about applications that can be used in brick-and-mortar stores using multilingual translation systems. Please also give specific examples, especially those using smart glasses."
[0149] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0150] Step 1:
[0151] The user puts on the smart glasses and uses the smart glasses' camera to take a picture of a label or guide sign written in the foreign language to be translated.
[0152] Input: Captured image (JPEG or PNG format)
[0153] Output: Raw image data for processing within the smart glasses
[0154] What it does: The camera in the smart glasses captures an image and stores it in the device's local memory.
[0155] Step 2:
[0156] The device converts the captured image into grayscale and performs preprocessing such as noise removal.
[0157] Input: Raw image data
[0158] Output: Preprocessed grayscale image
[0159] Specific operations: Using image processing means (OpenCV), convert the image to grayscale and apply a noise reduction filter.
[0160] Step 3:
[0161] The device extracts character regions from the preprocessed image and performs optical character recognition (OCR) to generate string data.
[0162] Input: Preprocessed grayscale image
[0163] Output: Extracted string data
[0164] Specific operation: Using character recognition means (Tesseract OCR), characters are detected in the image and the corresponding string of characters is extracted as digital text.
[0165] Step 4:
[0166] The server receives the extracted string data and translates it into the specified target language.
[0167] Input: String data (e.g., "Hello"), target language (e.g., "en")
[0168] Output: Translated text data (e.g. "Hello")
[0169] Specific operation: Using language translation means (GoogleTrans API), the extracted string is translated into the target language.
[0170] Step 5:
[0171] The server sends the translation results to the smart glasses display and displays them to the user.
[0172] Input: Translated text data
[0173] Output: Translation results displayed on the smart glasses display
[0174] Specific operation: Using the translation result output means, the translated text is displayed on the display of the smart glasses.
[0175] keyword
[0176] Generative AI model, prompt sentence
[0177] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0178] The present invention combines a multilingual translation system with an emotion engine. The system includes an image processing unit, a character recognition unit, a language translation unit, a translation result output unit, and an emotion engine.
[0179] First, the user inputs the manga image file (e.g., "manga_page.jpg") and the target language (e.g., English "en") into the device. The emotion engine then analyzes the user's voice input and facial images to recognize emotions.
[0180] The device then sends these input data to the server, with the request including the image file path, target language, and user emotion data.
[0181] The server analyzes the request received from the terminal, reads the image file using image processing means, and distinguishes between pictures and text. Specifically, the server converts the image file to grayscale and performs preprocessing to remove noise. Then, it uses character recognition means to extract text from the preprocessed image using optical character recognition (OCR) technology. This extracted text is temporarily stored.
[0182] The server then uses a language translation tool to translate the extracted characters into the target language. Generative AI is used to create a translation that is context-appropriate. Furthermore, the emotion engine adjusts the translation results based on the user's emotions as recognized. For example, if the user has positive emotions, the translation results will be adjusted to better reflect those emotions.
[0183] The translation result is passed to the translation result output means and converted into a data format for display on the user interface. The server then generates a response including the translation result and sends it to the terminal. The terminal receives the response from the server and displays the translation result to the user through the user interface.
[0184] As a concrete example, if a user translates a Japanese manga page "manga_page.jpg" into English and the emotion engine simultaneously recognizes the user's smiling face, the process will be as follows: When the user enters the image file path and the target language "en", the device will send this information and emotion data to the server. The server processes the image, extracts the string "hello" and translates it as "Hello". Because the emotion engine recognized the user's positive emotion, the translation result will be adjusted to "Hello! How are you feeling today?". This result will be sent to the device and displayed to the user.
[0185] In this way, the system of the present invention can provide high-quality, fast translations, while also achieving more natural and familiar translation results that match the user's emotions.
[0186] The processing flow will be explained below.
[0187] Step 1:
[0188] The user launches the application on their device and inputs the image file path of the manga they want to translate (e.g., "manga_page.jpg") and the target language (e.g., English "en"). The user also provides emotion data by speaking into the device or pointing their face at the camera.
[0189] Step 2:
[0190] The device receives the user's input data (image file path and target language) and simultaneously launches the emotion engine to analyze the user's emotions. The emotion engine recognizes the user's emotions by analyzing voice input and facial images.
[0191] Step 3:
[0192] The device sends the recognized emotion data (e.g., smile recognition result) along with the image file path and target language as a request to the server.
[0193] Step 4:
[0194] The server receives the request from the device and analyzes the image file path, target language, and emotion data.
[0195] Step 5:
[0196] The server starts the image processing function and reads the specified image file (e.g., "manga_page.jpg") using OpenCV. It converts the image to grayscale and performs preprocessing to remove noise.
[0197] Step 6:
[0198] The server distinguishes between pictures and text from the preprocessed images and identifies text regions.
[0199] Step 7:
[0200] The server uses a character recognition means to extract characters from the preprocessed image using optical character recognition (OCR) technology, for example, the string "hello" is extracted.
[0201] Step 8:
[0202] The server temporarily stores the extracted character data.
[0203] Step 9:
[0204] The server uses a language translation means to translate the extracted characters into the target language using generative artificial intelligence, for example, "hello" is translated into "hello".
[0205] Step 10:
[0206] The server uses the user's emotional data recognized by the emotion engine to adjust the expression of the translation result. For example, if the user is smiling, it adjusts the expression to a more positive one, such as "Hello! How are you feeling today?"
[0207] Step 11:
[0208] The adjusted translation result is passed to the translation result output means and converted into a data format for display on the user interface.
[0209] Step 12:
[0210] The server generates a response containing the final translation result and sends it to the terminal.
[0211] Step 13:
[0212] The device receives the response from the server and displays the translation result (e.g., "Hello! How are you feeling today?") to the user through a user interface, allowing the user to check the high-quality, emotion-based translation result.
[0213] Example 2
[0214] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0215] Conventional multilingual translation systems lack accuracy and naturalness in translation, making it difficult to provide translation results that reflect diverse emotions. Furthermore, the accuracy and speed of character extraction from image files are low, making it impossible to achieve translations that take into account the user's specific emotions. To solve these problems and quickly provide high-quality, natural translation results, a system that integrates image processing, character recognition, language translation, and emotion recognition technologies is needed.
[0216] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0217] In this invention, the server includes image processing means, character recognition means, language translation means, translation result output means, emotion recognition means, and user interface means. This enables the server to distinguish between pictures and text from image files, extract text using optical character recognition technology, and perform translation using a generative AI model. Furthermore, the server analyzes the user's emotions and adjusts the translation results based on those emotions, providing more natural and user-friendly translation results. The translation results are displayed through the user interface means, allowing the user to intuitively confirm the results.
[0218] "Image processing means" refers to the technology and devices used to distinguish between pictures and text in image files.
[0219] "Character recognition means" refers to a technique and device that extracts characters from an image using optical character recognition technology.
[0220] "Language translation tools" means technologies and devices that convert sentences or words into a different language, particularly those that use generative AI models to perform translation.
[0221] The "translation result output means" refers to a technology and device that converts the translation result into a data format for display on a user interface and outputs it.
[0222] "Emotion recognition means" refers to technology and devices that analyze the user's voice input and facial images to recognize emotions and adjust the translation results based on those emotions.
[0223] "User interface means" refers to the techniques and devices that allow a user to input data into the system and visually display the translation results.
[0224] The present invention combines a multilingual translation system with an emotion engine. This system includes image processing means, character recognition means, language translation means, translation result output means, and emotion recognition means. Specific implementations of these means are described below.
[0225] First, the user inputs an image file (e.g., "manga_page.jpg") and the target language into the device. In addition, the user inputs emotions into the system through voice input or facial images. These data are sent from the device to the server. The server analyzes the image file path, target language, and emotion data.
[0226] The server uses an image processing means to read the image file, first convert it to grayscale and remove noise, then process it to distinguish between pictures and text. Next, a character recognition means uses optical character recognition (OCR) technology to extract text from the preprocessed image. This technology can be implemented using common OCR software (e.g., Google OCR).
[0227] The extracted strings are temporarily stored on a server and then translated into the target language by a language translation tool using a generative AI model (e.g., OpenAI's GPT model) to provide highly accurate translation based on context.
[0228] Furthermore, the emotion recognition means analyzes the user's emotion data and adjusts the translation result based on the emotion. For example, if the user expresses a positive emotion, the translation result is adjusted to a more emotional expression.
[0229] Finally, the translation result is passed to the translation result output means, where it is converted into a data format suitable for the user interface. The server sends this response to the terminal, and the terminal displays the translation result to the user through the user interface.
[0230] As a concrete example, if a user translates a Japanese manga page "manga_page.jpg" into the target language, English, and simultaneously inputs a smiley face as emotion data, the process will be as follows: When the user inputs the image file path and the target language "en," the device sends this information and emotion data to the server. The server processes the image, extracts the string "hello" and translates it to "Hello." Because the emotion recognition means identifies the user's positive emotion, the translation result is adjusted to "Hello! How are you feeling today?" This result is sent to the device and displayed to the user.
[0231] Examples of prompts include:
[0232] Analyze the manga image file "manga_page.jpg". Then, translate the Japanese text extracted from this image into English, providing a translation that reflects a positive emotion because the user is smiling. The output should be in the following format: "Hello! How are you feeling today?".
[0233] In this way, the system of the present invention coordinates advanced processes, quickly provides high-quality translation results, and realizes natural, familiar translation that responds to the user's emotions.
[0234] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0235] Step 1:
[0236] The user inputs an image file and a target language into the terminal. Specifically, the user selects an image file called "manga_page.jpg" and English (en) as the target language. The user then inputs their own emotion (e.g., a smile) using a camera or microphone. The input data includes the image file path, the target language, and emotion data.
[0237] Input: Image file "manga_page.jpg", target language "en", emotion data (e.g., smile)
[0238] Output: Request data sent by the device to the server
[0239] Step 2:
[0240] The device sends the input data to the server. Specifically, the device combines the image file path "manga_page.jpg," the target language "en," and the emotion data into a single request data and sends it to the server.
[0241] Input: Request data (image file path, target language, emotion data)
[0242] Output: The request data received by the server
[0243] Step 3:
[0244] The server parses the request data. Specifically, the server receives the request data and separates and obtains the image file path, target language, and emotion data. Parsing is done using methods such as JSON parsing.
[0245] Input: Request data
[0246] Output: Individual data after analysis (image file path, target language, emotion data)
[0247] Step 4:
[0248] The server performs image processing. Specifically, the server reads the image file using image processing means, converts it to grayscale, then removes noise, and then performs processing to distinguish between pictures and text.
[0249] Input: Image file "manga_page.jpg"
[0250] Output: Image data after grayscale conversion and noise removal
[0251] Step 5:
[0252] The server extracts characters using optical character recognition (OCR) technology. Specifically, it extracts characters from the preprocessed image data using OCR technology. This process extracts the string "hello."
[0253] Input: Image data after grayscale conversion and noise removal
[0254] Output: The extracted string "Hello"
[0255] Step 6:
[0256] The server translates the string using a language translation tool. Specifically, it translates the extracted string "hello" into English using a generative AI model (e.g., GPT model). This translation results in the string "Hello."
[0257] Input: Extracted string "Hello"
[0258] Output: The translated string "Hello"
[0259] Step 7:
[0260] The server uses emotion recognition to adjust the translation result. Specifically, it adjusts the translation result based on emotion data (information that the user is smiling). As a result, the translated string "Hello" is changed to "Hello! How are you feeling today?"
[0261] Input: translated string "Hello", emotion data (e.g., smile)
[0262] Output: Adjusted translation result "Hello! How are you feeling today?"
[0263] Step 8:
[0264] The server converts the translation results into an output format. Specifically, it converts the adjusted translation results into a data format (e.g., HTML format) for display in a user interface.
[0265] Input: Adjusted translation result "Hello! How are you feeling today?"
[0266] Output: Data format for user interface display
[0267] Step 9:
[0268] The server generates response data and transmits it to the terminal. Specifically, the server generates response data including the converted translation result and transmits it to the terminal.
[0269] Input: Data format for user interface display
[0270] Output: Response data sent to the device
[0271] Step 10:
[0272] The device receives the response data from the server and displays the translation result to the user. Specifically, the device analyzes the response data received from the server and displays the translation result "Hello! How are you feeling today?" to the user through the user interface.
[0273] Input: Response data from the server
[0274] Output: The translation displayed to the user: "Hello! How are you feeling today?"
[0275] (Application example 2)
[0276] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0277] The present invention aims to improve the quality of communication in a multilingual translation system by not only reducing language barriers between employees but also by providing translation results that take into consideration the feelings of each employee. Another aim is to improve work efficiency by providing appropriate feedback when employees are tired or stressed.
[0278] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0279] In this invention, the server includes image processing means, character recognition means, language translation means, emotion recognition means, and emotion-based translation result adjustment means, which allows for extracting characters from images, translating them into a target language, and detecting the employee's emotional state and adjusting the translation result accordingly.
[0280] "Image processing means" refers to devices or software that have the function of extracting specific information from image files or analyzing images.
[0281] "Character recognition means" means a device or software that has the capability to use optical character recognition technology to detect characters in an image file and convert them into digital text.
[0282] "Language translation means" refers to a device or software capable of translating extracted digital text into another language.
[0283] "Translation result output means" refers to a device or software that has the function of providing an interface for displaying translated text to a user.
[0284] "Emotion recognition means" refers to a device or software that has the function of analyzing and recognizing emotions from a user's facial image or voice data.
[0285] "Means for adjusting translation results based on emotions" refers to devices or software that have the function of adjusting translation results based on recognized emotional data to provide more appropriate and natural expressions.
[0286] The present invention provides a system for supporting multilingual communication among employees in a factory and providing appropriate feedback based on the emotions of the employees. The system includes image processing means, character recognition means, language translation means, translation result output means, emotion recognition means, and emotion-based translation result adjustment means.
[0287] First, the user (employee) inputs the image file of the work instruction (e.g., "work_instruction.jpg") into the terminal. The emotion recognition means analyzes the user's facial image (e.g., "employee_face.jpg") to recognize the emotion. This data is sent to the server via the terminal.
[0288] The server first reads the image file using the image processing means, converts it to grayscale, and removes noise. Next, the character recognition means performs optical character recognition (OCR) to extract characters from the image. This extracted string is temporarily saved and translated into the target language (e.g., English "en") by the language translation means.
[0289] The emotion recognition means analyzes the user's facial image and obtains emotion data (e.g., tired, positive, negative). The emotion-based translation result adjustment means uses this emotion data to adjust the translation result to include appropriate feedback. For example, if the user is recognized as tired, the translation result is adjusted to say something like, "Next, please press this button. Is everything okay?"
[0290] Finally, the adjusted translation result is converted into a display format by the translation result output means, returned to the terminal, and displayed to the user.
[0291] As a concrete example, when a user translates a Japanese work instruction document "work_instruction.jpg" into English and the emotion engine simultaneously recognizes the user's tired expression, the process is as follows: When the user inputs the image file path and the target language "en", the device sends this information and emotion data to the server. The server processes the image, extracts the string "For the next step, please press this button," and translates it as "Next, please press this button." Because the emotion engine recognizes the user's tired expression, the translation result is adjusted to "Next, please press this button. Is everything okay?" This result is sent to the device and displayed to the user.
[0292] An example of an input prompt for a generative AI model is, "Extract Japanese text from the image file 'work_instruction.jpg' and translate it into English. In doing so, output a translation result that includes appropriate feedback based on the employee's emotions read from 'employee_face.jpg'."
[0293] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0294] Step 1:
[0295] The user inputs the image file of the work instruction (e.g., "work_instruction.jpg") and the user's face image (e.g., "employee_face.jpg") into the terminal. At this time, the target language is also input. The input data includes the image file path, the face image file path, and the target language information.
[0296] Step 2:
[0297] The terminal sends the input data (image file path, face image file path, target language) to the server. The server receives this data and passes it to the image processing means and emotion recognition means.
[0298] Step 3:
[0299] The server uses image processing means to read the image file ("work_instruction.jpg"), convert it to grayscale, and remove noise. The input data is the image file, and the output data is the preprocessed image file.
[0300] Step 4:
[0301] The server uses character recognition means to extract characters from the preprocessed image file. Specifically, it converts the characters into digital text using optical character recognition (OCR) technology. The input data is the preprocessed image file, and the output data is the extracted character string.
[0302] Step 5:
[0303] The server uses a language translation mechanism to translate the extracted string into the target language (e.g., English "en"), using generative artificial intelligence to provide a context-appropriate translation. The input data is the extracted string, and the output data is the translated string.
[0304] Step 6:
[0305] The server uses emotion recognition means to analyze the user's facial image ("employee_face.jpg") and recognize the user's emotional state. The input data is a facial image file, and the output data is recognized emotional data.
[0306] Step 7:
[0307] The server adjusts the translation result based on the recognized emotion data using a means for adjusting the translation result based on emotion. For example, if the server recognizes that the user is tired, it adds thoughtful feedback to the translation result. The input data is the translated string and emotion data, and the output data is the adjusted translation result.
[0308] Step 8:
[0309] The adjusted translation result is converted into a display format by the translation result output means and sent to the terminal. The terminal receives this data and displays it to the user. The input data is the adjusted translation result, and the output data is in a display format that the user can see.
[0310] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0311] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0312] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0313] [Second embodiment]
[0314] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0315] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0316] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0317] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0318] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0319] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0320] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0321] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0322] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0323] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0324] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0325] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0326] The present invention provides a multilingual translation system, which includes an image processing unit, a character recognition unit, a language translation unit, and a translation result output unit.
[0327] First, the user inputs the manga image file (e.g., "manga_page.jpg") and the target language (e.g., English "en") into the terminal. The terminal receives these inputs and sends the image file path and the target language to the server.
[0328] Next, the server receives the request from the terminal, reads the image file using image processing means, and distinguishes between pictures and text. Specifically, the server converts the image file to grayscale and performs preprocessing to remove noise.
[0329] The server then uses a character recognition tool to extract characters from the processed image. Optical character recognition technology is used to convert the characters in the image into digital text data. This extracted string of characters is temporarily stored.
[0330] The server then uses a language translation tool to translate the extracted characters into the target language. Generative artificial intelligence is used to convert the extracted text into the target language. This procedure allows for a more natural and context-appropriate translation.
[0331] The translated text is finally displayed on the user interface using the translation result output means. The terminal receives a response from the server and displays the translation result to the user. This process allows the user to obtain a fast and accurate translation result.
[0332] As a concrete example, suppose a user wants to translate a Japanese manga page "manga_page.jpg" into English. When the user enters the image file path and the target language "en" into the device, the device sends this information to the server. The server processes the image, extracts the string "Hello" and translates it to "Hello". Finally, the device displays the translation result "Hello" to the user.
[0333] This allows overseas manga fans to quickly enjoy a wide range of works in high quality, while domestic publishers can easily expand globally at low cost. Because the system is tailored to specific genres and styles, it is possible to provide appropriate translations that fit the context of each work.
[0334] The processing flow will be explained below.
[0335] Step 1:
[0336] The user launches the application on their device and enters the image file path of the manga they want to translate (e.g., "manga_page.jpg") and the target language (e.g., English "en").
[0337] Step 2:
[0338] The terminal receives the user's input data and sends it to the server as a request, which includes the image file path and the target language.
[0339] Step 3:
[0340] The server receives the request from the device and analyzes the image file path and target language.
[0341] Step 4:
[0342] The server starts the image processing means and reads the specified image file (e.g., "manga_page.jpg"). It converts the image file to grayscale using OpenCV and performs preprocessing to remove noise.
[0343] Step 5:
[0344] The server distinguishes between pictures and text from the preprocessed images and identifies areas containing text.
[0345] Step 6:
[0346] The server uses a character recognition means to extract characters from the pre-processed image using optical character recognition (OCR) techniques, for example, the text "hello" is extracted.
[0347] Step 7:
[0348] The server temporarily stores the extracted character data.
[0349] Step 8:
[0350] The server uses a language translation means to translate the extracted text into the target language using generative artificial intelligence, for example, "hello" is translated into "hello".
[0351] Step 9:
[0352] The server passes the translation result to the translation result output means.
[0353] Step 10:
[0354] The translation result output means formats the translation result and converts it into a data format for display on a user interface.
[0355] Step 11:
[0356] The server generates a response including the translation result and sends it to the terminal.
[0357] Step 12:
[0358] The terminal receives the response from the server and displays the translation result (e.g., "Hello") to the user through the user interface.
[0359] Example 1
[0360] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0361] Conventional multilingual translation systems have had difficulty properly recognizing text in images and providing natural, context-appropriate translations. Furthermore, building systems that enable end users to receive fast, accurate translation results is complex and expensive.
[0362] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0363] In this invention, the server includes an image processing means, a character recognition means, and a language translation means. This makes it possible to distinguish between pictures and text from image files, extract text using optical character recognition technology, and translate the extracted text into a target language using generative artificial intelligence. Specifically, by linking a terminal where a user inputs an image file and the target language with a server that receives and processes information from the terminal, a multilingual translation system that provides fast and accurate translation results is realized.
[0364] The "image processing means" is a means for reading an image file, distinguishing between pictures and text, and performing preprocessing such as grayscale conversion and noise removal as necessary.
[0365] "Character recognition means" refers to means that uses optical character recognition technology to extract characters from an image and convert them into digital text data.
[0366] The "language translation means" is a means for translating extracted text data into a target language using generative artificial intelligence.
[0367] The "translation result output means" is a means for displaying the translated text on a user interface.
[0368] A "terminal" is a device through which a user inputs an image file and a target language and sends that information to a server.
[0369] A "server" is a computer system that performs image processing, character recognition, and language translation based on information received from a terminal, and then sends the results to the terminal.
[0370] The present invention provides a multilingual translation system, which includes an image processing unit, a character recognition unit, a language translation unit, and a translation result output unit, and further includes a terminal where a user inputs an image file and a target language, and a server that processes information received from the terminal.
[0371] First, the user inputs the manga image file (e.g., "manga_page.jpg") and the target language (e.g., English "en") into the terminal. The terminal receives these inputs and sends the image file path and the target language to the server. The image file is usually a JPEG or PNG file, and the target language is based on the user's preference.
[0372] Next, the server receives the request from the device and uses image processing to read the image file and distinguish between pictures and text. Specifically, the server uses an image processing library such as OpenCV to convert the image file to grayscale and apply a noise reduction filter such as Gaussian blur. This preprocessing improves the accuracy of character recognition.
[0373] The server then uses character recognition to extract characters from the processed image. Optical character recognition software, such as Tesseract OCR, is used to convert the characters in the image into digital text data. This extracted string is then temporarily stored.
[0374] The server then uses a language translation tool to translate the extracted characters into the target language. Generative artificial intelligence (AI) is used to convert the extracted text into the target language. This procedure allows for a more natural and contextual translation. For example, a prompt sentence such as "Please translate the following Japanese text into English: 'Hello'" is sent to the generative AI model, and the translation result is "Hello."
[0375] Finally, the translated text is displayed on the user interface using the translation result output means. The terminal receives a response from the server and displays the translation result to the user. This process allows the user to obtain a fast and accurate translation result. This system allows, for example, overseas manga fans to enjoy Japanese manga in English, and domestic publishers to easily expand globally at low cost.
[0376] As a concrete example, if a user wants to translate a manga page "manga_page.jpg" into English, the user inputs the image file path and the target language "en" into their device. The device sends this information to the server, which performs image preprocessing, character recognition, and translation. As a result, the translated text "Hello" is displayed on the user's device. By utilizing a generative AI model, this system can provide natural-sounding translations that fit the context.
[0377] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0378] Step 1:
[0379] The user inputs the manga image file and target language into the device. Specifically, the user uses the device interface to select the path of the image file (e.g., "manga_page.jpg") and specify the target language (e.g., "en"). This inputs the image file path and target language information into the device.
[0380] Input: Image file (manga_page.jpg), target language (en)
[0381] Output: Image file path and target language information
[0382] Step 2:
[0383] The device sends the input information to the server. The device packages the image file path and target language information as an HTTP request and sends it to the server. Specifically, the information is included in the HTTP request body in JSON format or multipart form data.
[0384] Input: Image file path and target language information
[0385] Output: HTTP request sent to the server
[0386] Step 3:
[0387] The server preprocesses the image file. The server reads the image file received from the device, converts the image to grayscale using OpenCV, and removes noise using Gaussian blur, etc. This improves the accuracy of character recognition.
[0388] Input: HTTP request (image file path, target language information)
[0389] Output: Preprocessed image
[0390] Step 4:
[0391] The server performs character recognition. The server passes the preprocessed image to optical character recognition software such as Tesseract OCR, which extracts characters from the image. Specifically, it converts a string of characters, such as "hello," into digital text data and temporarily stores it.
[0392] Input: Preprocessed image
[0393] Output: Extracted text data (e.g. "Hello")
[0394] Step 5:
[0395] The server translates the characters into the target language. The server uses generative artificial intelligence (AI) to convert the extracted text data into the target language. The specific prompt is "Please translate the following Japanese text into English: 'Hello'". The translation result is "Hello".
[0396] Input: Extracted text data (e.g. "Hello")
[0397] Output: Translated text data (e.g. "Hello")
[0398] Step 6:
[0399] The server sends the translation results to the device. The server packages the translated text data in JSON format and sends it to the device as an HTTP response.
[0400] Input: Translated text data (e.g. "Hello")
[0401] Output: HTTP response sent to the device
[0402] Step 7:
[0403] The terminal displays the translation result to the user. The terminal analyzes the HTTP response received from the server and displays the translation result in the user interface. Specifically, it uses a text view or an alert box to allow the user to check the translation result (e.g., "Hello").
[0404] Input: HTTP response (translated text data)
[0405] Output: The translation result displayed in the user interface (e.g. "Hello")
[0406] (Application example 1)
[0407] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0408] In today's brick-and-mortar stores, store clerks need to be proficient in multiple languages to provide multilingual support, but there are limitations to this. In particular, in areas with a large number of tourists and foreign residents, fast and accurate translation is required, but this is difficult to do manually. Therefore, there is a need for a system that allows store clerks to instantly understand product labels and information signs written in foreign languages.
[0409] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0410] In this invention, the server includes an image processing unit, a character recognition unit, and a language translation unit. This allows images captured using the smart glasses to be translated in real time and the results to be visually displayed. This allows store clerks to quickly and accurately understand foreign language information, dramatically improving multilingual support.
[0411] "Image processing means" refers to a device or device function that reads an image file and performs preprocessing such as grayscale conversion and noise removal.
[0412] "Character recognition means" refers to a function that converts characters in an image into digital text data using optical character recognition (OCR) technology.
[0413] "Language translation means" refers to a device or software that has the function of translating extracted character data into a target language.
[0414] "Translation result output means" refers to a device or a function of a device for displaying translated text data to a user.
[0415] "Smart glasses" refers to a wearable device that has a built-in camera and display and has the ability to display information in real time.
[0416] "Means for real-time translation of images captured by smart glasses" refers to a device or software that has the function of instantly processing images captured by smart glasses and displaying the translation results.
[0417] This invention is a system that performs real-time multilingual translation in a physical store using smart glasses worn by store clerks. Specific embodiments of this system are described below.
[0418] System Overview
[0419] The system includes an image processing unit, a character recognition unit, a language translation unit, and a translation result output unit, as well as smart glasses and a unit for translating images captured by the glasses in real time.
[0420] Basic operation
[0421] The server analyzes the image entered by the user and generates the translation result using image processing, character recognition, and language translation. This series of processes uses the following hardware and software:
[0422] 1. Image processing means:
[0423] Hardware: Camera built into smart glasses
[0424] Software: OpenCV (image grayscale conversion and noise reduction)
[0425] 2. Character recognition means:
[0426] Software: Tesseract OCR (Optical Character Recognition Technology)
[0427] 3. Language Translation Methods:
[0428] Software: GoogleTrans API (language translation)
[0429] 4. Translation result output method:
[0430] Hardware: Smart glasses display
[0431] Software: Custom application for displaying translation results
[0432] Examples of data processing and data calculation
[0433] When a user (store clerk) wears smart glasses and sees a product label written in a foreign language, they take a picture of the label with their camera. The captured image is processed by the smart glasses' processor, which removes noise and converts it to grayscale. Next, Tesseract OCR is used to convert the characters in the image into digital text. This text data is then translated into the desired language using the Google Translate API. The final translation result is visually displayed on the smart glasses' display, allowing the user to instantly understand the content.
[0434] Specific examples
[0435] For example, suppose a user wears smart glasses and looks at the manga page "manga_page.jpg". The camera in the smart glasses captures this image and processes it based on the target language specified by the user (e.g., English "en"). OCR technology extracts the string "Hello", which is then translated into "Hello" by the GoogleTrans API. Finally, the translation result "Hello" is displayed on the display of the smart glasses.
[0436] Prompt Sentence Examples
[0437] "Please tell us more about applications that can be used in brick-and-mortar stores using multilingual translation systems. Please also give specific examples, especially those using smart glasses."
[0438] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0439] Step 1:
[0440] The user puts on the smart glasses and uses the smart glasses' camera to take a picture of a label or guide sign written in the foreign language to be translated.
[0441] Input: Captured image (JPEG or PNG format)
[0442] Output: Raw image data for processing within the smart glasses
[0443] What it does: The camera in the smart glasses captures an image and stores it in the device's local memory.
[0444] Step 2:
[0445] The device converts the captured image into grayscale and performs preprocessing such as noise removal.
[0446] Input: Raw image data
[0447] Output: Preprocessed grayscale image
[0448] Specific operations: Using image processing means (OpenCV), convert the image to grayscale and apply a noise reduction filter.
[0449] Step 3:
[0450] The device extracts character regions from the preprocessed image and performs optical character recognition (OCR) to generate string data.
[0451] Input: Preprocessed grayscale image
[0452] Output: Extracted string data
[0453] Specific operation: Using character recognition means (Tesseract OCR), characters are detected in the image and the corresponding string of characters is extracted as digital text.
[0454] Step 4:
[0455] The server receives the extracted string data and translates it into the specified target language.
[0456] Input: String data (e.g., "Hello"), target language (e.g., "en")
[0457] Output: Translated text data (e.g. "Hello")
[0458] Specific operation: Using language translation means (GoogleTrans API), the extracted string is translated into the target language.
[0459] Step 5:
[0460] The server sends the translation results to the smart glasses display and displays them to the user.
[0461] Input: Translated text data
[0462] Output: Translation results displayed on the smart glasses display
[0463] Specific operation: Using the translation result output means, the translated text is displayed on the display of the smart glasses.
[0464] keyword
[0465] Generative AI model, prompt sentence
[0466] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0467] The present invention combines a multilingual translation system with an emotion engine. The system includes an image processing unit, a character recognition unit, a language translation unit, a translation result output unit, and an emotion engine.
[0468] First, the user inputs the manga image file (e.g., "manga_page.jpg") and the target language (e.g., English "en") into the device. The emotion engine then analyzes the user's voice input and facial images to recognize emotions.
[0469] The device then sends these input data to the server, with the request including the image file path, target language, and user emotion data.
[0470] The server analyzes the request received from the terminal, reads the image file using image processing means, and distinguishes between pictures and text. Specifically, the server converts the image file to grayscale and performs preprocessing to remove noise. Then, it uses character recognition means to extract text from the preprocessed image using optical character recognition (OCR) technology. This extracted text is temporarily stored.
[0471] The server then uses a language translation tool to translate the extracted characters into the target language. Generative AI is used to create a translation that is context-appropriate. Furthermore, the emotion engine adjusts the translation results based on the user's emotions as recognized. For example, if the user has positive emotions, the translation results will be adjusted to better reflect those emotions.
[0472] The translation result is passed to the translation result output means and converted into a data format for display on the user interface. The server then generates a response including the translation result and sends it to the terminal. The terminal receives the response from the server and displays the translation result to the user through the user interface.
[0473] As a concrete example, if a user translates a Japanese manga page "manga_page.jpg" into English and the emotion engine simultaneously recognizes the user's smiling face, the process will be as follows: When the user enters the image file path and the target language "en", the device will send this information and emotion data to the server. The server processes the image, extracts the string "hello" and translates it as "Hello". Because the emotion engine recognized the user's positive emotion, the translation result will be adjusted to "Hello! How are you feeling today?". This result will be sent to the device and displayed to the user.
[0474] In this way, the system of the present invention can provide high-quality, fast translations, while also achieving more natural and familiar translation results that match the user's emotions.
[0475] The processing flow will be explained below.
[0476] Step 1:
[0477] The user launches the application on their device and inputs the image file path of the manga they want to translate (e.g., "manga_page.jpg") and the target language (e.g., English "en"). The user also provides emotion data by speaking into the device or pointing their face at the camera.
[0478] Step 2:
[0479] The device receives the user's input data (image file path and target language) and simultaneously launches the emotion engine to analyze the user's emotions. The emotion engine recognizes the user's emotions by analyzing voice input and facial images.
[0480] Step 3:
[0481] The device sends the recognized emotion data (e.g., smile recognition result) along with the image file path and target language as a request to the server.
[0482] Step 4:
[0483] The server receives the request from the device and analyzes the image file path, target language, and emotion data.
[0484] Step 5:
[0485] The server starts the image processing function and reads the specified image file (e.g., "manga_page.jpg") using OpenCV. It converts the image to grayscale and performs preprocessing to remove noise.
[0486] Step 6:
[0487] The server distinguishes between pictures and text from the preprocessed images and identifies text regions.
[0488] Step 7:
[0489] The server uses a character recognition means to extract characters from the preprocessed image using optical character recognition (OCR) technology, for example, the string "hello" is extracted.
[0490] Step 8:
[0491] The server temporarily stores the extracted character data.
[0492] Step 9:
[0493] The server uses a language translation means to translate the extracted characters into the target language using generative artificial intelligence, for example, "hello" is translated into "hello".
[0494] Step 10:
[0495] The server uses the user's emotional data recognized by the emotion engine to adjust the expression of the translation result. For example, if the user is smiling, it adjusts the expression to a more positive one, such as "Hello! How are you feeling today?"
[0496] Step 11:
[0497] The adjusted translation result is passed to the translation result output means and converted into a data format for display on the user interface.
[0498] Step 12:
[0499] The server generates a response containing the final translation result and sends it to the terminal.
[0500] Step 13:
[0501] The device receives the response from the server and displays the translation result (e.g., "Hello! How are you feeling today?") to the user through a user interface, allowing the user to check the high-quality, emotion-based translation result.
[0502] Example 2
[0503] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0504] Conventional multilingual translation systems lack accuracy and naturalness in translation, making it difficult to provide translation results that reflect diverse emotions. Furthermore, the accuracy and speed of character extraction from image files are low, making it impossible to achieve translations that take into account the user's specific emotions. To solve these problems and quickly provide high-quality, natural translation results, a system that integrates image processing, character recognition, language translation, and emotion recognition technologies is needed.
[0505] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0506] In this invention, the server includes image processing means, character recognition means, language translation means, translation result output means, emotion recognition means, and user interface means. This enables the server to distinguish between pictures and text from image files, extract text using optical character recognition technology, and perform translation using a generative AI model. Furthermore, the server analyzes the user's emotions and adjusts the translation results based on those emotions, providing more natural and user-friendly translation results. The translation results are displayed through the user interface means, allowing the user to intuitively confirm the results.
[0507] "Image processing means" refers to the technology and devices used to distinguish between pictures and text in image files.
[0508] "Character recognition means" refers to a technique and device that extracts characters from an image using optical character recognition technology.
[0509] "Language translation tools" means technologies and devices that convert sentences or words into a different language, particularly those that use generative AI models to perform translation.
[0510] The "translation result output means" refers to a technology and device that converts the translation result into a data format for display on a user interface and outputs it.
[0511] "Emotion recognition means" refers to technology and devices that analyze the user's voice input and facial images to recognize emotions and adjust the translation results based on those emotions.
[0512] "User interface means" refers to the techniques and devices that allow a user to input data into the system and visually display the translation results.
[0513] The present invention combines a multilingual translation system with an emotion engine. This system includes image processing means, character recognition means, language translation means, translation result output means, and emotion recognition means. Specific implementations of these means are described below.
[0514] First, the user inputs an image file (e.g., "manga_page.jpg") and the target language into the device. In addition, the user inputs emotions into the system through voice input or facial images. These data are sent from the device to the server. The server analyzes the image file path, target language, and emotion data.
[0515] The server uses an image processing means to read the image file, first convert it to grayscale and remove noise, then process it to distinguish between pictures and text. Next, a character recognition means uses optical character recognition (OCR) technology to extract text from the preprocessed image. This technology can be implemented using common OCR software (e.g., Google OCR).
[0516] The extracted strings are temporarily stored on a server and then translated into the target language by a language translation tool using a generative AI model (e.g., OpenAI's GPT model) to provide highly accurate translation based on context.
[0517] Furthermore, the emotion recognition means analyzes the user's emotion data and adjusts the translation result based on the emotion. For example, if the user expresses a positive emotion, the translation result is adjusted to a more emotional expression.
[0518] Finally, the translation result is passed to the translation result output means, where it is converted into a data format suitable for the user interface. The server sends this response to the terminal, and the terminal displays the translation result to the user through the user interface.
[0519] As a concrete example, if a user translates a Japanese manga page "manga_page.jpg" into the target language, English, and simultaneously inputs a smiley face as emotion data, the process will be as follows: When the user inputs the image file path and the target language "en," the device sends this information and emotion data to the server. The server processes the image, extracts the string "hello" and translates it to "Hello." Because the emotion recognition means identifies the user's positive emotion, the translation result is adjusted to "Hello! How are you feeling today?" This result is sent to the device and displayed to the user.
[0520] Examples of prompts include:
[0521] Analyze the manga image file "manga_page.jpg". Then, translate the Japanese text extracted from this image into English, providing a translation that reflects a positive emotion because the user is smiling. The output should be in the following format: "Hello! How are you feeling today?".
[0522] In this way, the system of the present invention coordinates advanced processes, quickly provides high-quality translation results, and realizes natural, familiar translation that responds to the user's emotions.
[0523] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0524] Step 1:
[0525] The user inputs an image file and a target language into the terminal. Specifically, the user selects an image file called "manga_page.jpg" and English (en) as the target language. The user then inputs their own emotion (e.g., a smile) using a camera or microphone. The input data includes the image file path, the target language, and emotion data.
[0526] Input: Image file "manga_page.jpg", target language "en", emotion data (e.g., smile)
[0527] Output: Request data sent by the device to the server
[0528] Step 2:
[0529] The device sends the input data to the server. Specifically, the device combines the image file path "manga_page.jpg," the target language "en," and the emotion data into a single request data and sends it to the server.
[0530] Input: Request data (image file path, target language, emotion data)
[0531] Output: The request data received by the server
[0532] Step 3:
[0533] The server parses the request data. Specifically, the server receives the request data and separates and obtains the image file path, target language, and emotion data. Parsing is done using methods such as JSON parsing.
[0534] Input: Request data
[0535] Output: Individual data after analysis (image file path, target language, emotion data)
[0536] Step 4:
[0537] The server performs image processing. Specifically, the server reads the image file using image processing means, converts it to grayscale, then removes noise, and then performs processing to distinguish between pictures and text.
[0538] Input: Image file "manga_page.jpg"
[0539] Output: Image data after grayscale conversion and noise removal
[0540] Step 5:
[0541] The server extracts characters using optical character recognition (OCR) technology. Specifically, it extracts characters from the preprocessed image data using OCR technology. This process extracts the string "hello."
[0542] Input: Image data after grayscale conversion and noise removal
[0543] Output: The extracted string "Hello"
[0544] Step 6:
[0545] The server translates the string using a language translation tool. Specifically, it translates the extracted string "hello" into English using a generative AI model (e.g., GPT model). This translation results in the string "Hello."
[0546] Input: Extracted string "Hello"
[0547] Output: The translated string "Hello"
[0548] Step 7:
[0549] The server uses emotion recognition to adjust the translation result. Specifically, it adjusts the translation result based on emotion data (information that the user is smiling). As a result, the translated string "Hello" is changed to "Hello! How are you feeling today?"
[0550] Input: translated string "Hello", emotion data (e.g., smile)
[0551] Output: Adjusted translation result "Hello! How are you feeling today?"
[0552] Step 8:
[0553] The server converts the translation results into an output format. Specifically, it converts the adjusted translation results into a data format (e.g., HTML format) for display in a user interface.
[0554] Input: Adjusted translation result "Hello! How are you feeling today?"
[0555] Output: Data format for user interface display
[0556] Step 9:
[0557] The server generates response data and transmits it to the terminal. Specifically, the server generates response data including the converted translation result and transmits it to the terminal.
[0558] Input: Data format for user interface display
[0559] Output: Response data sent to the device
[0560] Step 10:
[0561] The device receives the response data from the server and displays the translation result to the user. Specifically, the device analyzes the response data received from the server and displays the translation result "Hello! How are you feeling today?" to the user through the user interface.
[0562] Input: Response data from the server
[0563] Output: The translation displayed to the user: "Hello! How are you feeling today?"
[0564] (Application example 2)
[0565] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0566] The present invention aims to improve the quality of communication in a multilingual translation system by not only reducing language barriers between employees but also by providing translation results that take into consideration the feelings of each employee. Another aim is to improve work efficiency by providing appropriate feedback when employees are tired or stressed.
[0567] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0568] In this invention, the server includes image processing means, character recognition means, language translation means, emotion recognition means, and emotion-based translation result adjustment means, which allows for extracting characters from images, translating them into a target language, and detecting the employee's emotional state and adjusting the translation result accordingly.
[0569] "Image processing means" refers to devices or software that have the function of extracting specific information from image files or analyzing images.
[0570] "Character recognition means" means a device or software that has the capability to use optical character recognition technology to detect characters in an image file and convert them into digital text.
[0571] "Language translation means" refers to a device or software capable of translating extracted digital text into another language.
[0572] "Translation result output means" refers to a device or software that has the function of providing an interface for displaying translated text to a user.
[0573] "Emotion recognition means" refers to a device or software that has the function of analyzing and recognizing emotions from a user's facial image or voice data.
[0574] "Means for adjusting translation results based on emotions" refers to devices or software that have the function of adjusting translation results based on recognized emotional data to provide more appropriate and natural expressions.
[0575] The present invention provides a system for supporting multilingual communication among employees in a factory and providing appropriate feedback based on the emotions of the employees. The system includes image processing means, character recognition means, language translation means, translation result output means, emotion recognition means, and emotion-based translation result adjustment means.
[0576] First, the user (employee) inputs the image file of the work instruction (e.g., "work_instruction.jpg") into the terminal. The emotion recognition means analyzes the user's facial image (e.g., "employee_face.jpg") to recognize the emotion. This data is sent to the server via the terminal.
[0577] The server first reads the image file using the image processing means, converts it to grayscale, and removes noise. Next, the character recognition means performs optical character recognition (OCR) to extract characters from the image. This extracted string is temporarily saved and translated into the target language (e.g., English "en") by the language translation means.
[0578] The emotion recognition means analyzes the user's facial image and obtains emotion data (e.g., tired, positive, negative). The emotion-based translation result adjustment means uses this emotion data to adjust the translation result to include appropriate feedback. For example, if the user is recognized as tired, the translation result is adjusted to say something like, "Next, please press this button. Is everything okay?"
[0579] Finally, the adjusted translation result is converted into a display format by the translation result output means, returned to the terminal, and displayed to the user.
[0580] As a concrete example, when a user translates a Japanese work instruction document "work_instruction.jpg" into English and the emotion engine simultaneously recognizes the user's tired expression, the process is as follows: When the user inputs the image file path and the target language "en", the device sends this information and emotion data to the server. The server processes the image, extracts the string "For the next step, please press this button," and translates it as "Next, please press this button." Because the emotion engine recognizes the user's tired expression, the translation result is adjusted to "Next, please press this button. Is everything okay?" This result is sent to the device and displayed to the user.
[0581] An example of an input prompt for a generative AI model is, "Extract Japanese text from the image file 'work_instruction.jpg' and translate it into English. In doing so, output a translation result that includes appropriate feedback based on the employee's emotions read from 'employee_face.jpg'."
[0582] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0583] Step 1:
[0584] The user inputs the image file of the work instruction (e.g., "work_instruction.jpg") and the user's face image (e.g., "employee_face.jpg") into the terminal. At this time, the target language is also input. The input data includes the image file path, the face image file path, and the target language information.
[0585] Step 2:
[0586] The terminal sends the input data (image file path, face image file path, target language) to the server. The server receives this data and passes it to the image processing means and emotion recognition means.
[0587] Step 3:
[0588] The server uses image processing means to read the image file ("work_instruction.jpg"), convert it to grayscale, and remove noise. The input data is the image file, and the output data is the preprocessed image file.
[0589] Step 4:
[0590] The server uses character recognition means to extract characters from the preprocessed image file. Specifically, it converts the characters into digital text using optical character recognition (OCR) technology. The input data is the preprocessed image file, and the output data is the extracted character string.
[0591] Step 5:
[0592] The server uses a language translation mechanism to translate the extracted string into the target language (e.g., English "en"), using generative artificial intelligence to provide a context-appropriate translation. The input data is the extracted string, and the output data is the translated string.
[0593] Step 6:
[0594] The server uses emotion recognition means to analyze the user's facial image ("employee_face.jpg") and recognize the user's emotional state. The input data is a facial image file, and the output data is recognized emotional data.
[0595] Step 7:
[0596] The server adjusts the translation result based on the recognized emotion data using a means for adjusting the translation result based on emotion. For example, if the server recognizes that the user is tired, it adds thoughtful feedback to the translation result. The input data is the translated string and emotion data, and the output data is the adjusted translation result.
[0597] Step 8:
[0598] The adjusted translation result is converted into a display format by the translation result output means and sent to the terminal. The terminal receives this data and displays it to the user. The input data is the adjusted translation result, and the output data is in a display format that the user can see.
[0599] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0600] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0601] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0602] [Third embodiment]
[0603] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0604] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0605] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0606] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0607] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0608] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0609] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0610] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0611] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0612] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0613] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0614] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0615] The present invention provides a multilingual translation system, which includes an image processing unit, a character recognition unit, a language translation unit, and a translation result output unit.
[0616] First, the user inputs the manga image file (e.g., "manga_page.jpg") and the target language (e.g., English "en") into the terminal. The terminal receives these inputs and sends the image file path and the target language to the server.
[0617] Next, the server receives the request from the terminal, reads the image file using image processing means, and distinguishes between pictures and text. Specifically, the server converts the image file to grayscale and performs preprocessing to remove noise.
[0618] The server then uses a character recognition tool to extract characters from the processed image. Optical character recognition technology is used to convert the characters in the image into digital text data. This extracted string of characters is temporarily stored.
[0619] The server then uses a language translation tool to translate the extracted characters into the target language. Generative artificial intelligence is used to convert the extracted text into the target language. This procedure allows for a more natural and context-appropriate translation.
[0620] The translated text is finally displayed on the user interface using the translation result output means. The terminal receives a response from the server and displays the translation result to the user. This process allows the user to obtain a fast and accurate translation result.
[0621] As a concrete example, suppose a user wants to translate a Japanese manga page "manga_page.jpg" into English. When the user enters the image file path and the target language "en" into the device, the device sends this information to the server. The server processes the image, extracts the string "Hello" and translates it to "Hello". Finally, the device displays the translation result "Hello" to the user.
[0622] This allows overseas manga fans to quickly enjoy a wide range of works in high quality, while domestic publishers can easily expand globally at low cost. Because the system is tailored to specific genres and styles, it is possible to provide appropriate translations that fit the context of each work.
[0623] The processing flow will be explained below.
[0624] Step 1:
[0625] The user launches the application on their device and enters the image file path of the manga they want to translate (e.g., "manga_page.jpg") and the target language (e.g., English "en").
[0626] Step 2:
[0627] The terminal receives the user's input data and sends it to the server as a request, which includes the image file path and the target language.
[0628] Step 3:
[0629] The server receives the request from the device and analyzes the image file path and target language.
[0630] Step 4:
[0631] The server starts the image processing means and reads the specified image file (e.g., "manga_page.jpg"). It converts the image file to grayscale using OpenCV and performs preprocessing to remove noise.
[0632] Step 5:
[0633] The server distinguishes between pictures and text from the preprocessed images and identifies areas containing text.
[0634] Step 6:
[0635] The server uses a character recognition means to extract characters from the pre-processed image using optical character recognition (OCR) techniques, for example, the text "hello" is extracted.
[0636] Step 7:
[0637] The server temporarily stores the extracted character data.
[0638] Step 8:
[0639] The server uses a language translation means to translate the extracted text into the target language using generative artificial intelligence, for example, "hello" is translated into "hello".
[0640] Step 9:
[0641] The server passes the translation result to the translation result output means.
[0642] Step 10:
[0643] The translation result output means formats the translation result and converts it into a data format for display on a user interface.
[0644] Step 11:
[0645] The server generates a response including the translation result and sends it to the terminal.
[0646] Step 12:
[0647] The terminal receives the response from the server and displays the translation result (e.g., "Hello") to the user through the user interface.
[0648] Example 1
[0649] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0650] Conventional multilingual translation systems have had difficulty properly recognizing text in images and providing natural, context-appropriate translations. Furthermore, building systems that enable end users to receive fast, accurate translation results is complex and expensive.
[0651] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0652] In this invention, the server includes an image processing means, a character recognition means, and a language translation means. This makes it possible to distinguish between pictures and text from image files, extract text using optical character recognition technology, and translate the extracted text into a target language using generative artificial intelligence. Specifically, by linking a terminal where a user inputs an image file and the target language with a server that receives and processes information from the terminal, a multilingual translation system that provides fast and accurate translation results is realized.
[0653] The "image processing means" is a means for reading an image file, distinguishing between pictures and text, and performing preprocessing such as grayscale conversion and noise removal as necessary.
[0654] "Character recognition means" refers to means that uses optical character recognition technology to extract characters from an image and convert them into digital text data.
[0655] The "language translation means" is a means for translating extracted text data into a target language using generative artificial intelligence.
[0656] The "translation result output means" is a means for displaying the translated text on a user interface.
[0657] A "terminal" is a device through which a user inputs an image file and a target language and sends that information to a server.
[0658] A "server" is a computer system that performs image processing, character recognition, and language translation based on information received from a terminal, and then sends the results to the terminal.
[0659] The present invention provides a multilingual translation system, which includes an image processing unit, a character recognition unit, a language translation unit, and a translation result output unit, and further includes a terminal where a user inputs an image file and a target language, and a server that processes information received from the terminal.
[0660] First, the user inputs the manga image file (e.g., "manga_page.jpg") and the target language (e.g., English "en") into the terminal. The terminal receives these inputs and sends the image file path and the target language to the server. The image file is usually a JPEG or PNG file, and the target language is based on the user's preference.
[0661] Next, the server receives the request from the device and uses image processing to read the image file and distinguish between pictures and text. Specifically, the server uses an image processing library such as OpenCV to convert the image file to grayscale and apply a noise reduction filter such as Gaussian blur. This preprocessing improves the accuracy of character recognition.
[0662] The server then uses character recognition to extract characters from the processed image. Optical character recognition software, such as Tesseract OCR, is used to convert the characters in the image into digital text data. This extracted string is then temporarily stored.
[0663] The server then uses a language translation tool to translate the extracted characters into the target language. Generative artificial intelligence (AI) is used to convert the extracted text into the target language. This procedure allows for a more natural and contextual translation. For example, a prompt sentence such as "Please translate the following Japanese text into English: 'Hello'" is sent to the generative AI model, and the translation result is "Hello."
[0664] Finally, the translated text is displayed on the user interface using the translation result output means. The terminal receives a response from the server and displays the translation result to the user. This process allows the user to obtain a fast and accurate translation result. This system allows, for example, overseas manga fans to enjoy Japanese manga in English, and domestic publishers to easily expand globally at low cost.
[0665] As a concrete example, if a user wants to translate a manga page "manga_page.jpg" into English, the user inputs the image file path and the target language "en" into their device. The device sends this information to the server, which performs image preprocessing, character recognition, and translation. As a result, the translated text "Hello" is displayed on the user's device. By utilizing a generative AI model, this system can provide natural-sounding translations that fit the context.
[0666] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0667] Step 1:
[0668] The user inputs the manga image file and target language into the device. Specifically, the user uses the device interface to select the path of the image file (e.g., "manga_page.jpg") and specify the target language (e.g., "en"). This inputs the image file path and target language information into the device.
[0669] Input: Image file (manga_page.jpg), target language (en)
[0670] Output: Image file path and target language information
[0671] Step 2:
[0672] The device sends the input information to the server. The device packages the image file path and target language information as an HTTP request and sends it to the server. Specifically, the information is included in the HTTP request body in JSON format or multipart form data.
[0673] Input: Image file path and target language information
[0674] Output: HTTP request sent to the server
[0675] Step 3:
[0676] The server preprocesses the image file. The server reads the image file received from the device, converts the image to grayscale using OpenCV, and removes noise using Gaussian blur, etc. This improves the accuracy of character recognition.
[0677] Input: HTTP request (image file path, target language information)
[0678] Output: Preprocessed image
[0679] Step 4:
[0680] The server performs character recognition. The server passes the preprocessed image to optical character recognition software such as Tesseract OCR, which extracts characters from the image. Specifically, it converts a string of characters, such as "hello," into digital text data and temporarily stores it.
[0681] Input: Preprocessed image
[0682] Output: Extracted text data (e.g. "Hello")
[0683] Step 5:
[0684] The server translates the characters into the target language. The server uses generative artificial intelligence (AI) to convert the extracted text data into the target language. The specific prompt is "Please translate the following Japanese text into English: 'Hello'". The translation result is "Hello".
[0685] Input: Extracted text data (e.g. "Hello")
[0686] Output: Translated text data (e.g. "Hello")
[0687] Step 6:
[0688] The server sends the translation results to the device. The server packages the translated text data in JSON format and sends it to the device as an HTTP response.
[0689] Input: Translated text data (e.g. "Hello")
[0690] Output: HTTP response sent to the device
[0691] Step 7:
[0692] The terminal displays the translation result to the user. The terminal analyzes the HTTP response received from the server and displays the translation result in the user interface. Specifically, it uses a text view or an alert box to allow the user to check the translation result (e.g., "Hello").
[0693] Input: HTTP response (translated text data)
[0694] Output: The translation result displayed in the user interface (e.g. "Hello")
[0695] (Application example 1)
[0696] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0697] In today's brick-and-mortar stores, store clerks need to be proficient in multiple languages to provide multilingual support, but there are limitations to this. In particular, in areas with a large number of tourists and foreign residents, fast and accurate translation is required, but this is difficult to do manually. Therefore, there is a need for a system that allows store clerks to instantly understand product labels and information signs written in foreign languages.
[0698] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0699] In this invention, the server includes an image processing unit, a character recognition unit, and a language translation unit. This allows images captured using the smart glasses to be translated in real time and the results to be visually displayed. This allows store clerks to quickly and accurately understand foreign language information, dramatically improving multilingual support.
[0700] "Image processing means" refers to a device or device function that reads an image file and performs preprocessing such as grayscale conversion and noise removal.
[0701] "Character recognition means" refers to a function that converts characters in an image into digital text data using optical character recognition (OCR) technology.
[0702] "Language translation means" refers to a device or software that has the function of translating extracted character data into a target language.
[0703] "Translation result output means" refers to a device or a function of a device for displaying translated text data to a user.
[0704] "Smart glasses" refers to a wearable device that has a built-in camera and display and has the ability to display information in real time.
[0705] "Means for real-time translation of images captured by smart glasses" refers to a device or software that has the function of instantly processing images captured by smart glasses and displaying the translation results.
[0706] This invention is a system that performs real-time multilingual translation in a physical store using smart glasses worn by store clerks. Specific embodiments of this system are described below.
[0707] System Overview
[0708] The system includes an image processing unit, a character recognition unit, a language translation unit, and a translation result output unit, as well as smart glasses and a unit for translating images captured by the glasses in real time.
[0709] Basic operation
[0710] The server analyzes the image entered by the user and generates the translation result using image processing, character recognition, and language translation. This series of processes uses the following hardware and software:
[0711] 1. Image processing means:
[0712] Hardware: Camera built into smart glasses
[0713] Software: OpenCV (image grayscale conversion and noise reduction)
[0714] 2. Character recognition means:
[0715] Software: Tesseract OCR (Optical Character Recognition Technology)
[0716] 3. Language Translation Methods:
[0717] Software: GoogleTrans API (language translation)
[0718] 4. Translation result output method:
[0719] Hardware: Smart glasses display
[0720] Software: Custom application for displaying translation results
[0721] Examples of data processing and data calculation
[0722] When a user (store clerk) wears smart glasses and sees a product label written in a foreign language, they take a picture of the label with their camera. The captured image is processed by the smart glasses' processor, which removes noise and converts it to grayscale. Next, Tesseract OCR is used to convert the characters in the image into digital text. This text data is then translated into the desired language using the Google Translate API. The final translation result is visually displayed on the smart glasses' display, allowing the user to instantly understand the content.
[0723] Specific examples
[0724] For example, suppose a user wears smart glasses and looks at the manga page "manga_page.jpg". The camera in the smart glasses captures this image and processes it based on the target language specified by the user (e.g., English "en"). OCR technology extracts the string "Hello", which is then translated into "Hello" by the GoogleTrans API. Finally, the translation result "Hello" is displayed on the display of the smart glasses.
[0725] Prompt Sentence Examples
[0726] "Please tell us more about applications that can be used in brick-and-mortar stores using multilingual translation systems. Please also give specific examples, especially those using smart glasses."
[0727] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0728] Step 1:
[0729] The user puts on the smart glasses and uses the smart glasses' camera to take a picture of a label or guide sign written in the foreign language to be translated.
[0730] Input: Captured image (JPEG or PNG format)
[0731] Output: Raw image data for processing within the smart glasses
[0732] What it does: The camera in the smart glasses captures an image and stores it in the device's local memory.
[0733] Step 2:
[0734] The device converts the captured image into grayscale and performs preprocessing such as noise removal.
[0735] Input: Raw image data
[0736] Output: Preprocessed grayscale image
[0737] Specific operations: Using image processing means (OpenCV), convert the image to grayscale and apply a noise reduction filter.
[0738] Step 3:
[0739] The device extracts character regions from the preprocessed image and performs optical character recognition (OCR) to generate string data.
[0740] Input: Preprocessed grayscale image
[0741] Output: Extracted string data
[0742] Specific operation: Using character recognition means (Tesseract OCR), characters are detected in the image and the corresponding string of characters is extracted as digital text.
[0743] Step 4:
[0744] The server receives the extracted string data and translates it into the specified target language.
[0745] Input: String data (e.g., "Hello"), target language (e.g., "en")
[0746] Output: Translated text data (e.g. "Hello")
[0747] Specific operation: Using language translation means (GoogleTrans API), the extracted string is translated into the target language.
[0748] Step 5:
[0749] The server sends the translation results to the smart glasses display and displays them to the user.
[0750] Input: Translated text data
[0751] Output: Translation results displayed on the smart glasses display
[0752] Specific operation: Using the translation result output means, the translated text is displayed on the display of the smart glasses.
[0753] keyword
[0754] Generative AI model, prompt sentence
[0755] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0756] The present invention combines a multilingual translation system with an emotion engine. The system includes an image processing unit, a character recognition unit, a language translation unit, a translation result output unit, and an emotion engine.
[0757] First, the user inputs the manga image file (e.g., "manga_page.jpg") and the target language (e.g., English "en") into the device. The emotion engine then analyzes the user's voice input and facial images to recognize emotions.
[0758] The device then sends these input data to the server, with the request including the image file path, target language, and user emotion data.
[0759] The server analyzes the request received from the terminal, reads the image file using image processing means, and distinguishes between pictures and text. Specifically, the server converts the image file to grayscale and performs preprocessing to remove noise. Then, it uses character recognition means to extract text from the preprocessed image using optical character recognition (OCR) technology. This extracted text is temporarily stored.
[0760] The server then uses a language translation tool to translate the extracted characters into the target language. Generative AI is used to create a translation that is context-appropriate. Furthermore, the emotion engine adjusts the translation results based on the user's emotions as recognized. For example, if the user has positive emotions, the translation results will be adjusted to better reflect those emotions.
[0761] The translation result is passed to the translation result output means and converted into a data format for display on the user interface. The server then generates a response including the translation result and sends it to the terminal. The terminal receives the response from the server and displays the translation result to the user through the user interface.
[0762] As a concrete example, if a user translates a Japanese manga page "manga_page.jpg" into English and the emotion engine simultaneously recognizes the user's smiling face, the process will be as follows: When the user enters the image file path and the target language "en", the device will send this information and emotion data to the server. The server processes the image, extracts the string "hello" and translates it as "Hello". Because the emotion engine recognized the user's positive emotion, the translation result will be adjusted to "Hello! How are you feeling today?". This result will be sent to the device and displayed to the user.
[0763] In this way, the system of the present invention can provide high-quality, fast translations, while also achieving more natural and familiar translation results that match the user's emotions.
[0764] The processing flow will be explained below.
[0765] Step 1:
[0766] The user launches the application on their device and inputs the image file path of the manga they want to translate (e.g., "manga_page.jpg") and the target language (e.g., English "en"). The user also provides emotion data by speaking into the device or pointing their face at the camera.
[0767] Step 2:
[0768] The device receives the user's input data (image file path and target language) and simultaneously launches the emotion engine to analyze the user's emotions. The emotion engine recognizes the user's emotions by analyzing voice input and facial images.
[0769] Step 3:
[0770] The device sends the recognized emotion data (e.g., smile recognition result) along with the image file path and target language as a request to the server.
[0771] Step 4:
[0772] The server receives the request from the device and analyzes the image file path, target language, and emotion data.
[0773] Step 5:
[0774] The server starts the image processing function and reads the specified image file (e.g., "manga_page.jpg") using OpenCV. It converts the image to grayscale and performs preprocessing to remove noise.
[0775] Step 6:
[0776] The server distinguishes between pictures and text from the preprocessed images and identifies text regions.
[0777] Step 7:
[0778] The server uses a character recognition means to extract characters from the preprocessed image using optical character recognition (OCR) technology, for example, the string "hello" is extracted.
[0779] Step 8:
[0780] The server temporarily stores the extracted character data.
[0781] Step 9:
[0782] The server uses a language translation means to translate the extracted characters into the target language using generative artificial intelligence, for example, "hello" is translated into "hello".
[0783] Step 10:
[0784] The server uses the user's emotional data recognized by the emotion engine to adjust the expression of the translation result. For example, if the user is smiling, it adjusts the expression to a more positive one, such as "Hello! How are you feeling today?"
[0785] Step 11:
[0786] The adjusted translation result is passed to the translation result output means and converted into a data format for display on the user interface.
[0787] Step 12:
[0788] The server generates a response containing the final translation result and sends it to the terminal.
[0789] Step 13:
[0790] The device receives the response from the server and displays the translation result (e.g., "Hello! How are you feeling today?") to the user through a user interface, allowing the user to check the high-quality, emotion-based translation result.
[0791] Example 2
[0792] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0793] Conventional multilingual translation systems lack accuracy and naturalness in translation, making it difficult to provide translation results that reflect diverse emotions. Furthermore, the accuracy and speed of character extraction from image files are low, making it impossible to achieve translations that take into account the user's specific emotions. To solve these problems and quickly provide high-quality, natural translation results, a system that integrates image processing, character recognition, language translation, and emotion recognition technologies is needed.
[0794] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0795] In this invention, the server includes image processing means, character recognition means, language translation means, translation result output means, emotion recognition means, and user interface means. This enables the server to distinguish between pictures and text from image files, extract text using optical character recognition technology, and perform translation using a generative AI model. Furthermore, the server analyzes the user's emotions and adjusts the translation results based on those emotions, providing more natural and user-friendly translation results. The translation results are displayed through the user interface means, allowing the user to intuitively confirm the results.
[0796] "Image processing means" refers to the technology and devices used to distinguish between pictures and text in image files.
[0797] "Character recognition means" refers to a technique and device that extracts characters from an image using optical character recognition technology.
[0798] "Language translation tools" means technologies and devices that convert sentences or words into a different language, particularly those that use generative AI models to perform translation.
[0799] The "translation result output means" refers to a technology and device that converts the translation result into a data format for display on a user interface and outputs it.
[0800] "Emotion recognition means" refers to technology and devices that analyze the user's voice input and facial images to recognize emotions and adjust the translation results based on those emotions.
[0801] "User interface means" refers to the techniques and devices that allow a user to input data into the system and visually display the translation results.
[0802] The present invention combines a multilingual translation system with an emotion engine. This system includes image processing means, character recognition means, language translation means, translation result output means, and emotion recognition means. Specific implementations of these means are described below.
[0803] First, the user inputs an image file (e.g., "manga_page.jpg") and the target language into the device. In addition, the user inputs emotions into the system through voice input or facial images. These data are sent from the device to the server. The server analyzes the image file path, target language, and emotion data.
[0804] The server uses an image processing means to read the image file, first convert it to grayscale and remove noise, then process it to distinguish between pictures and text. Next, a character recognition means uses optical character recognition (OCR) technology to extract text from the preprocessed image. This technology can be implemented using common OCR software (e.g., Google OCR).
[0805] The extracted strings are temporarily stored on a server and then translated into the target language by a language translation tool using a generative AI model (e.g., OpenAI's GPT model) to provide highly accurate translation based on context.
[0806] Furthermore, the emotion recognition means analyzes the user's emotion data and adjusts the translation result based on the emotion. For example, if the user expresses a positive emotion, the translation result is adjusted to a more emotional expression.
[0807] Finally, the translation result is passed to the translation result output means, where it is converted into a data format suitable for the user interface. The server sends this response to the terminal, and the terminal displays the translation result to the user through the user interface.
[0808] As a concrete example, if a user translates a Japanese manga page "manga_page.jpg" into the target language, English, and simultaneously inputs a smiley face as emotion data, the process will be as follows: When the user inputs the image file path and the target language "en," the device sends this information and emotion data to the server. The server processes the image, extracts the string "hello" and translates it to "Hello." Because the emotion recognition means identifies the user's positive emotion, the translation result is adjusted to "Hello! How are you feeling today?" This result is sent to the device and displayed to the user.
[0809] Examples of prompts include:
[0810] Analyze the manga image file "manga_page.jpg". Then, translate the Japanese text extracted from this image into English, providing a translation that reflects a positive emotion because the user is smiling. The output should be in the following format: "Hello! How are you feeling today?".
[0811] In this way, the system of the present invention coordinates advanced processes, quickly provides high-quality translation results, and realizes natural, familiar translation that responds to the user's emotions.
[0812] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0813] Step 1:
[0814] The user inputs an image file and a target language into the terminal. Specifically, the user selects an image file called "manga_page.jpg" and English (en) as the target language. The user then inputs their own emotion (e.g., a smile) using a camera or microphone. The input data includes the image file path, the target language, and emotion data.
[0815] Input: Image file "manga_page.jpg", target language "en", emotion data (e.g., smile)
[0816] Output: Request data sent by the device to the server
[0817] Step 2:
[0818] The device sends the input data to the server. Specifically, the device combines the image file path "manga_page.jpg," the target language "en," and the emotion data into a single request data and sends it to the server.
[0819] Input: Request data (image file path, target language, emotion data)
[0820] Output: The request data received by the server
[0821] Step 3:
[0822] The server parses the request data. Specifically, the server receives the request data and separates and obtains the image file path, target language, and emotion data. Parsing is done using methods such as JSON parsing.
[0823] Input: Request data
[0824] Output: Individual data after analysis (image file path, target language, emotion data)
[0825] Step 4:
[0826] The server performs image processing. Specifically, the server reads the image file using image processing means, converts it to grayscale, then removes noise, and then performs processing to distinguish between pictures and text.
[0827] Input: Image file "manga_page.jpg"
[0828] Output: Image data after grayscale conversion and noise removal
[0829] Step 5:
[0830] The server extracts characters using optical character recognition (OCR) technology. Specifically, it extracts characters from the preprocessed image data using OCR technology. This process extracts the string "hello."
[0831] Input: Image data after grayscale conversion and noise removal
[0832] Output: The extracted string "Hello"
[0833] Step 6:
[0834] The server translates the string using a language translation tool. Specifically, it translates the extracted string "hello" into English using a generative AI model (e.g., GPT model). This translation results in the string "Hello."
[0835] Input: Extracted string "Hello"
[0836] Output: The translated string "Hello"
[0837] Step 7:
[0838] The server uses emotion recognition to adjust the translation result. Specifically, it adjusts the translation result based on emotion data (information that the user is smiling). As a result, the translated string "Hello" is changed to "Hello! How are you feeling today?"
[0839] Input: translated string "Hello", emotion data (e.g., smile)
[0840] Output: Adjusted translation result "Hello! How are you feeling today?"
[0841] Step 8:
[0842] The server converts the translation results into an output format. Specifically, it converts the adjusted translation results into a data format (e.g., HTML format) for display in a user interface.
[0843] Input: Adjusted translation result "Hello! How are you feeling today?"
[0844] Output: Data format for user interface display
[0845] Step 9:
[0846] The server generates response data and transmits it to the terminal. Specifically, the server generates response data including the converted translation result and transmits it to the terminal.
[0847] Input: Data format for user interface display
[0848] Output: Response data sent to the device
[0849] Step 10:
[0850] The device receives the response data from the server and displays the translation result to the user. Specifically, the device analyzes the response data received from the server and displays the translation result "Hello! How are you feeling today?" to the user through the user interface.
[0851] Input: Response data from the server
[0852] Output: The translation displayed to the user: "Hello! How are you feeling today?"
[0853] (Application example 2)
[0854] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0855] The present invention aims to improve the quality of communication in a multilingual translation system by not only reducing language barriers between employees but also by providing translation results that take into consideration the feelings of each employee. Another aim is to improve work efficiency by providing appropriate feedback when employees are tired or stressed.
[0856] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0857] In this invention, the server includes image processing means, character recognition means, language translation means, emotion recognition means, and emotion-based translation result adjustment means, which allows for extracting characters from images, translating them into a target language, and detecting the employee's emotional state and adjusting the translation result accordingly.
[0858] "Image processing means" refers to devices or software that have the function of extracting specific information from image files or analyzing images.
[0859] "Character recognition means" means a device or software that has the capability to use optical character recognition technology to detect characters in an image file and convert them into digital text.
[0860] "Language translation means" refers to a device or software capable of translating extracted digital text into another language.
[0861] "Translation result output means" refers to a device or software that has the function of providing an interface for displaying translated text to a user.
[0862] "Emotion recognition means" refers to a device or software that has the function of analyzing and recognizing emotions from a user's facial image or voice data.
[0863] "Means for adjusting translation results based on emotions" refers to devices or software that have the function of adjusting translation results based on recognized emotional data to provide more appropriate and natural expressions.
[0864] The present invention provides a system for supporting multilingual communication among employees in a factory and providing appropriate feedback based on the emotions of the employees. The system includes image processing means, character recognition means, language translation means, translation result output means, emotion recognition means, and emotion-based translation result adjustment means.
[0865] First, the user (employee) inputs the image file of the work instruction (e.g., "work_instruction.jpg") into the terminal. The emotion recognition means analyzes the user's facial image (e.g., "employee_face.jpg") to recognize the emotion. This data is sent to the server via the terminal.
[0866] The server first reads the image file using the image processing means, converts it to grayscale, and removes noise. Next, the character recognition means performs optical character recognition (OCR) to extract characters from the image. This extracted string is temporarily saved and translated into the target language (e.g., English "en") by the language translation means.
[0867] The emotion recognition means analyzes the user's facial image and obtains emotion data (e.g., tired, positive, negative). The emotion-based translation result adjustment means uses this emotion data to adjust the translation result to include appropriate feedback. For example, if the user is recognized as tired, the translation result is adjusted to say something like, "Next, please press this button. Is everything okay?"
[0868] Finally, the adjusted translation result is converted into a display format by the translation result output means, returned to the terminal, and displayed to the user.
[0869] As a concrete example, when a user translates a Japanese work instruction document "work_instruction.jpg" into English and the emotion engine simultaneously recognizes the user's tired expression, the process is as follows: When the user inputs the image file path and the target language "en", the device sends this information and emotion data to the server. The server processes the image, extracts the string "For the next step, please press this button," and translates it as "Next, please press this button." Because the emotion engine recognizes the user's tired expression, the translation result is adjusted to "Next, please press this button. Is everything okay?" This result is sent to the device and displayed to the user.
[0870] An example of an input prompt for a generative AI model is, "Extract Japanese text from the image file 'work_instruction.jpg' and translate it into English. In doing so, output a translation result that includes appropriate feedback based on the employee's emotions read from 'employee_face.jpg'."
[0871] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0872] Step 1:
[0873] The user inputs the image file of the work instruction (e.g., "work_instruction.jpg") and the user's face image (e.g., "employee_face.jpg") into the terminal. At this time, the target language is also input. The input data includes the image file path, the face image file path, and the target language information.
[0874] Step 2:
[0875] The terminal sends the input data (image file path, face image file path, target language) to the server. The server receives this data and passes it to the image processing means and emotion recognition means.
[0876] Step 3:
[0877] The server uses image processing means to read the image file ("work_instruction.jpg"), convert it to grayscale, and remove noise. The input data is the image file, and the output data is the preprocessed image file.
[0878] Step 4:
[0879] The server uses character recognition means to extract characters from the preprocessed image file. Specifically, it converts the characters into digital text using optical character recognition (OCR) technology. The input data is the preprocessed image file, and the output data is the extracted character string.
[0880] Step 5:
[0881] The server uses a language translation mechanism to translate the extracted string into the target language (e.g., English "en"), using generative artificial intelligence to provide a context-appropriate translation. The input data is the extracted string, and the output data is the translated string.
[0882] Step 6:
[0883] The server uses emotion recognition means to analyze the user's facial image ("employee_face.jpg") and recognize the user's emotional state. The input data is a facial image file, and the output data is recognized emotional data.
[0884] Step 7:
[0885] The server adjusts the translation result based on the recognized emotion data using a means for adjusting the translation result based on emotion. For example, if the server recognizes that the user is tired, it adds thoughtful feedback to the translation result. The input data is the translated string and emotion data, and the output data is the adjusted translation result.
[0886] Step 8:
[0887] The adjusted translation result is converted into a display format by the translation result output means and sent to the terminal. The terminal receives this data and displays it to the user. The input data is the adjusted translation result, and the output data is in a display format that the user can see.
[0888] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0889] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0890] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0891] [Fourth embodiment]
[0892] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0893] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0894] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0895] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0896] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0897] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0898] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0899] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0900] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0901] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0902] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0903] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0904] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0905] The present invention provides a multilingual translation system, which includes an image processing unit, a character recognition unit, a language translation unit, and a translation result output unit.
[0906] First, the user inputs the manga image file (e.g., "manga_page.jpg") and the target language (e.g., English "en") into the terminal. The terminal receives these inputs and sends the image file path and the target language to the server.
[0907] Next, the server receives the request from the terminal, reads the image file using image processing means, and distinguishes between pictures and text. Specifically, the server converts the image file to grayscale and performs preprocessing to remove noise.
[0908] The server then uses a character recognition tool to extract characters from the processed image. Optical character recognition technology is used to convert the characters in the image into digital text data. This extracted string of characters is temporarily stored.
[0909] The server then uses a language translation tool to translate the extracted characters into the target language. Generative artificial intelligence is used to convert the extracted text into the target language. This procedure allows for a more natural and context-appropriate translation.
[0910] The translated text is finally displayed on the user interface using the translation result output means. The terminal receives a response from the server and displays the translation result to the user. This process allows the user to obtain a fast and accurate translation result.
[0911] As a concrete example, suppose a user wants to translate a Japanese manga page "manga_page.jpg" into English. When the user enters the image file path and the target language "en" into the device, the device sends this information to the server. The server processes the image, extracts the string "Hello" and translates it to "Hello". Finally, the device displays the translation result "Hello" to the user.
[0912] This allows overseas manga fans to quickly enjoy a wide range of works in high quality, while domestic publishers can easily expand globally at low cost. Because the system is tailored to specific genres and styles, it is possible to provide appropriate translations that fit the context of each work.
[0913] The processing flow will be explained below.
[0914] Step 1:
[0915] The user launches the application on their device and enters the image file path of the manga they want to translate (e.g., "manga_page.jpg") and the target language (e.g., English "en").
[0916] Step 2:
[0917] The terminal receives the user's input data and sends it to the server as a request, which includes the image file path and the target language.
[0918] Step 3:
[0919] The server receives the request from the device and analyzes the image file path and target language.
[0920] Step 4:
[0921] The server starts the image processing means and reads the specified image file (e.g., "manga_page.jpg"). It converts the image file to grayscale using OpenCV and performs preprocessing to remove noise.
[0922] Step 5:
[0923] The server distinguishes between pictures and text from the preprocessed images and identifies areas containing text.
[0924] Step 6:
[0925] The server uses a character recognition means to extract characters from the pre-processed image using optical character recognition (OCR) techniques, for example, the text "hello" is extracted.
[0926] Step 7:
[0927] The server temporarily stores the extracted character data.
[0928] Step 8:
[0929] The server uses a language translation means to translate the extracted text into the target language using generative artificial intelligence, for example, "hello" is translated into "hello".
[0930] Step 9:
[0931] The server passes the translation result to the translation result output means.
[0932] Step 10:
[0933] The translation result output means formats the translation result and converts it into a data format for display on a user interface.
[0934] Step 11:
[0935] The server generates a response including the translation result and sends it to the terminal.
[0936] Step 12:
[0937] The terminal receives the response from the server and displays the translation result (e.g., "Hello") to the user through the user interface.
[0938] Example 1
[0939] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0940] Conventional multilingual translation systems have had difficulty properly recognizing text in images and providing natural, context-appropriate translations. Furthermore, building systems that enable end users to receive fast, accurate translation results is complex and expensive.
[0941] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0942] In this invention, the server includes an image processing means, a character recognition means, and a language translation means. This makes it possible to distinguish between pictures and text from image files, extract text using optical character recognition technology, and translate the extracted text into a target language using generative artificial intelligence. Specifically, by linking a terminal where a user inputs an image file and the target language with a server that receives and processes information from the terminal, a multilingual translation system that provides fast and accurate translation results is realized.
[0943] The "image processing means" is a means for reading an image file, distinguishing between pictures and text, and performing preprocessing such as grayscale conversion and noise removal as necessary.
[0944] "Character recognition means" refers to means that uses optical character recognition technology to extract characters from an image and convert them into digital text data.
[0945] The "language translation means" is a means for translating extracted text data into a target language using generative artificial intelligence.
[0946] The "translation result output means" is a means for displaying the translated text on a user interface.
[0947] A "terminal" is a device through which a user inputs an image file and a target language and sends that information to a server.
[0948] A "server" is a computer system that performs image processing, character recognition, and language translation based on information received from a terminal, and then sends the results to the terminal.
[0949] The present invention provides a multilingual translation system, which includes an image processing unit, a character recognition unit, a language translation unit, and a translation result output unit, and further includes a terminal where a user inputs an image file and a target language, and a server that processes information received from the terminal.
[0950] First, the user inputs the manga image file (e.g., "manga_page.jpg") and the target language (e.g., English "en") into the terminal. The terminal receives these inputs and sends the image file path and the target language to the server. The image file is usually a JPEG or PNG file, and the target language is based on the user's preference.
[0951] Next, the server receives the request from the device and uses image processing to read the image file and distinguish between pictures and text. Specifically, the server uses an image processing library such as OpenCV to convert the image file to grayscale and apply a noise reduction filter such as Gaussian blur. This preprocessing improves the accuracy of character recognition.
[0952] The server then uses character recognition to extract characters from the processed image. Optical character recognition software, such as Tesseract OCR, is used to convert the characters in the image into digital text data. This extracted string is then temporarily stored.
[0953] The server then uses a language translation tool to translate the extracted characters into the target language. Generative artificial intelligence (AI) is used to convert the extracted text into the target language. This procedure allows for a more natural and contextual translation. For example, a prompt sentence such as "Please translate the following Japanese text into English: 'Hello'" is sent to the generative AI model, and the translation result is "Hello."
[0954] Finally, the translated text is displayed on the user interface using the translation result output means. The terminal receives a response from the server and displays the translation result to the user. This process allows the user to obtain a fast and accurate translation result. This system allows, for example, overseas manga fans to enjoy Japanese manga in English, and domestic publishers to easily expand globally at low cost.
[0955] As a concrete example, if a user wants to translate a manga page "manga_page.jpg" into English, the user inputs the image file path and the target language "en" into their device. The device sends this information to the server, which performs image preprocessing, character recognition, and translation. As a result, the translated text "Hello" is displayed on the user's device. By utilizing a generative AI model, this system can provide natural-sounding translations that fit the context.
[0956] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0957] Step 1:
[0958] The user inputs the manga image file and target language into the device. Specifically, the user uses the device interface to select the path of the image file (e.g., "manga_page.jpg") and specify the target language (e.g., "en"). This inputs the image file path and target language information into the device.
[0959] Input: Image file (manga_page.jpg), target language (en)
[0960] Output: Image file path and target language information
[0961] Step 2:
[0962] The device sends the input information to the server. The device packages the image file path and target language information as an HTTP request and sends it to the server. Specifically, the information is included in the HTTP request body in JSON format or multipart form data.
[0963] Input: Image file path and target language information
[0964] Output: HTTP request sent to the server
[0965] Step 3:
[0966] The server preprocesses the image file. The server reads the image file received from the device, converts the image to grayscale using OpenCV, and removes noise using Gaussian blur, etc. This improves the accuracy of character recognition.
[0967] Input: HTTP request (image file path, target language information)
[0968] Output: Preprocessed image
[0969] Step 4:
[0970] The server performs character recognition. The server passes the preprocessed image to optical character recognition software such as Tesseract OCR, which extracts characters from the image. Specifically, it converts a string of characters, such as "hello," into digital text data and temporarily stores it.
[0971] Input: Preprocessed image
[0972] Output: Extracted text data (e.g. "Hello")
[0973] Step 5:
[0974] The server translates the characters into the target language. The server uses generative artificial intelligence (AI) to convert the extracted text data into the target language. The specific prompt is "Please translate the following Japanese text into English: 'Hello'". The translation result is "Hello".
[0975] Input: Extracted text data (e.g. "Hello")
[0976] Output: Translated text data (e.g. "Hello")
[0977] Step 6:
[0978] The server sends the translation results to the device. The server packages the translated text data in JSON format and sends it to the device as an HTTP response.
[0979] Input: Translated text data (e.g. "Hello")
[0980] Output: HTTP response sent to the device
[0981] Step 7:
[0982] The terminal displays the translation result to the user. The terminal analyzes the HTTP response received from the server and displays the translation result in the user interface. Specifically, it uses a text view or an alert box to allow the user to check the translation result (e.g., "Hello").
[0983] Input: HTTP response (translated text data)
[0984] Output: The translation result displayed in the user interface (e.g. "Hello")
[0985] (Application example 1)
[0986] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0987] In today's brick-and-mortar stores, store clerks need to be proficient in multiple languages to provide multilingual support, but there are limitations to this. In particular, in areas with a large number of tourists and foreign residents, fast and accurate translation is required, but this is difficult to do manually. Therefore, there is a need for a system that allows store clerks to instantly understand product labels and information signs written in foreign languages.
[0988] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0989] In this invention, the server includes an image processing unit, a character recognition unit, and a language translation unit. This allows images captured using the smart glasses to be translated in real time and the results to be visually displayed. This allows store clerks to quickly and accurately understand foreign language information, dramatically improving multilingual support.
[0990] "Image processing means" refers to a device or device function that reads an image file and performs preprocessing such as grayscale conversion and noise removal.
[0991] "Character recognition means" refers to a function that converts characters in an image into digital text data using optical character recognition (OCR) technology.
[0992] "Language translation means" refers to a device or software that has the function of translating extracted character data into a target language.
[0993] "Translation result output means" refers to a device or a function of a device for displaying translated text data to a user.
[0994] "Smart glasses" refers to a wearable device that has a built-in camera and display and has the ability to display information in real time.
[0995] "Means for real-time translation of images captured by smart glasses" refers to a device or software that has the function of instantly processing images captured by smart glasses and displaying the translation results.
[0996] This invention is a system that performs real-time multilingual translation in a physical store using smart glasses worn by store clerks. Specific embodiments of this system are described below.
[0997] System Overview
[0998] The system includes an image processing unit, a character recognition unit, a language translation unit, and a translation result output unit, as well as smart glasses and a unit for translating images captured by the glasses in real time.
[0999] Basic operation
[1000] The server analyzes the image entered by the user and generates the translation result using image processing, character recognition, and language translation. This series of processes uses the following hardware and software:
[1001] 1. Image processing means:
[1002] Hardware: Camera built into smart glasses
[1003] Software: OpenCV (image grayscale conversion and noise reduction)
[1004] 2. Character recognition means:
[1005] Software: Tesseract OCR (Optical Character Recognition Technology)
[1006] 3. Language Translation Methods:
[1007] Software: GoogleTrans API (language translation)
[1008] 4. Translation result output method:
[1009] Hardware: Smart glasses display
[1010] Software: Custom application for displaying translation results
[1011] Examples of data processing and data calculation
[1012] When a user (store clerk) wears smart glasses and sees a product label written in a foreign language, they take a picture of the label with their camera. The captured image is processed by the smart glasses' processor, which removes noise and converts it to grayscale. Next, Tesseract OCR is used to convert the characters in the image into digital text. This text data is then translated into the desired language using the Google Translate API. The final translation result is visually displayed on the smart glasses' display, allowing the user to instantly understand the content.
[1013] Specific examples
[1014] For example, suppose a user wears smart glasses and looks at the manga page "manga_page.jpg". The camera in the smart glasses captures this image and processes it based on the target language specified by the user (e.g., English "en"). OCR technology extracts the string "Hello", which is then translated into "Hello" by the GoogleTrans API. Finally, the translation result "Hello" is displayed on the display of the smart glasses.
[1015] Prompt Sentence Examples
[1016] "Please tell us more about applications that can be used in brick-and-mortar stores using multilingual translation systems. Please also give specific examples, especially those using smart glasses."
[1017] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1018] Step 1:
[1019] The user puts on the smart glasses and uses the smart glasses' camera to take a picture of a label or guide sign written in the foreign language to be translated.
[1020] Input: Captured image (JPEG or PNG format)
[1021] Output: Raw image data for processing within the smart glasses
[1022] What it does: The camera in the smart glasses captures an image and stores it in the device's local memory.
[1023] Step 2:
[1024] The device converts the captured image into grayscale and performs preprocessing such as noise removal.
[1025] Input: Raw image data
[1026] Output: Preprocessed grayscale image
[1027] Specific operations: Using image processing means (OpenCV), convert the image to grayscale and apply a noise reduction filter.
[1028] Step 3:
[1029] The device extracts character regions from the preprocessed image and performs optical character recognition (OCR) to generate string data.
[1030] Input: Preprocessed grayscale image
[1031] Output: Extracted string data
[1032] Specific operation: Using character recognition means (Tesseract OCR), characters are detected in the image and the corresponding string of characters is extracted as digital text.
[1033] Step 4:
[1034] The server receives the extracted string data and translates it into the specified target language.
[1035] Input: String data (e.g., "Hello"), target language (e.g., "en")
[1036] Output: Translated text data (e.g. "Hello")
[1037] Specific operation: Using language translation means (GoogleTrans API), the extracted string is translated into the target language.
[1038] Step 5:
[1039] The server sends the translation results to the smart glasses display and displays them to the user.
[1040] Input: Translated text data
[1041] Output: Translation results displayed on the smart glasses display
[1042] Specific operation: Using the translation result output means, the translated text is displayed on the display of the smart glasses.
[1043] keyword
[1044] Generative AI model, prompt sentence
[1045] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1046] The present invention combines a multilingual translation system with an emotion engine. The system includes an image processing unit, a character recognition unit, a language translation unit, a translation result output unit, and an emotion engine.
[1047] First, the user inputs the manga image file (e.g., "manga_page.jpg") and the target language (e.g., English "en") into the device. The emotion engine then analyzes the user's voice input and facial images to recognize emotions.
[1048] The device then sends these input data to the server, with the request including the image file path, target language, and user emotion data.
[1049] The server analyzes the request received from the terminal, reads the image file using image processing means, and distinguishes between pictures and text. Specifically, the server converts the image file to grayscale and performs preprocessing to remove noise. Then, it uses character recognition means to extract text from the preprocessed image using optical character recognition (OCR) technology. This extracted text is temporarily stored.
[1050] The server then uses a language translation tool to translate the extracted characters into the target language. Generative AI is used to create a translation that is context-appropriate. Furthermore, the emotion engine adjusts the translation results based on the user's emotions as recognized. For example, if the user has positive emotions, the translation results will be adjusted to better reflect those emotions.
[1051] The translation result is passed to the translation result output means and converted into a data format for display on the user interface. The server then generates a response including the translation result and sends it to the terminal. The terminal receives the response from the server and displays the translation result to the user through the user interface.
[1052] As a concrete example, if a user translates a Japanese manga page "manga_page.jpg" into English and the emotion engine simultaneously recognizes the user's smiling face, the process will be as follows: When the user enters the image file path and the target language "en", the device will send this information and emotion data to the server. The server processes the image, extracts the string "hello" and translates it as "Hello". Because the emotion engine recognized the user's positive emotion, the translation result will be adjusted to "Hello! How are you feeling today?". This result will be sent to the device and displayed to the user.
[1053] In this way, the system of the present invention can provide high-quality, fast translations, while also achieving more natural and familiar translation results that match the user's emotions.
[1054] The processing flow will be explained below.
[1055] Step 1:
[1056] The user launches the application on their device and inputs the image file path of the manga they want to translate (e.g., "manga_page.jpg") and the target language (e.g., English "en"). The user also provides emotion data by speaking into the device or pointing their face at the camera.
[1057] Step 2:
[1058] The device receives the user's input data (image file path and target language) and simultaneously launches the emotion engine to analyze the user's emotions. The emotion engine recognizes the user's emotions by analyzing voice input and facial images.
[1059] Step 3:
[1060] The device sends the recognized emotion data (e.g., smile recognition result) along with the image file path and target language as a request to the server.
[1061] Step 4:
[1062] The server receives the request from the device and analyzes the image file path, target language, and emotion data.
[1063] Step 5:
[1064] The server starts the image processing function and reads the specified image file (e.g., "manga_page.jpg") using OpenCV. It converts the image to grayscale and performs preprocessing to remove noise.
[1065] Step 6:
[1066] The server distinguishes between pictures and text from the preprocessed images and identifies text regions.
[1067] Step 7:
[1068] The server uses a character recognition means to extract characters from the preprocessed image using optical character recognition (OCR) technology, for example, the string "hello" is extracted.
[1069] Step 8:
[1070] The server temporarily stores the extracted character data.
[1071] Step 9:
[1072] The server uses a language translation means to translate the extracted characters into the target language using generative artificial intelligence, for example, "hello" is translated into "hello".
[1073] Step 10:
[1074] The server uses the user's emotional data recognized by the emotion engine to adjust the expression of the translation result. For example, if the user is smiling, it adjusts the expression to a more positive one, such as "Hello! How are you feeling today?"
[1075] Step 11:
[1076] The adjusted translation result is passed to the translation result output means and converted into a data format for display on the user interface.
[1077] Step 12:
[1078] The server generates a response containing the final translation result and sends it to the terminal.
[1079] Step 13:
[1080] The device receives the response from the server and displays the translation result (e.g., "Hello! How are you feeling today?") to the user through a user interface, allowing the user to check the high-quality, emotion-based translation result.
[1081] Example 2
[1082] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1083] Conventional multilingual translation systems lack accuracy and naturalness in translation, making it difficult to provide translation results that reflect diverse emotions. Furthermore, the accuracy and speed of character extraction from image files are low, making it impossible to achieve translations that take into account the user's specific emotions. To solve these problems and quickly provide high-quality, natural translation results, a system that integrates image processing, character recognition, language translation, and emotion recognition technologies is needed.
[1084] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1085] In this invention, the server includes image processing means, character recognition means, language translation means, translation result output means, emotion recognition means, and user interface means. This enables the server to distinguish between pictures and text from image files, extract text using optical character recognition technology, and perform translation using a generative AI model. Furthermore, the server analyzes the user's emotions and adjusts the translation results based on those emotions, providing more natural and user-friendly translation results. The translation results are displayed through the user interface means, allowing the user to intuitively confirm the results.
[1086] "Image processing means" refers to the technology and devices used to distinguish between pictures and text in image files.
[1087] "Character recognition means" refers to a technique and device that extracts characters from an image using optical character recognition technology.
[1088] "Language translation tools" means technologies and devices that convert sentences or words into a different language, particularly those that use generative AI models to perform translation.
[1089] The "translation result output means" refers to a technology and device that converts the translation result into a data format for display on a user interface and outputs it.
[1090] "Emotion recognition means" refers to technology and devices that analyze the user's voice input and facial images to recognize emotions and adjust the translation results based on those emotions.
[1091] "User interface means" refers to the techniques and devices that allow a user to input data into the system and visually display the translation results.
[1092] The present invention combines a multilingual translation system with an emotion engine. This system includes image processing means, character recognition means, language translation means, translation result output means, and emotion recognition means. Specific implementations of these means are described below.
[1093] First, the user inputs an image file (e.g., "manga_page.jpg") and the target language into the device. In addition, the user inputs emotions into the system through voice input or facial images. These data are sent from the device to the server. The server analyzes the image file path, target language, and emotion data.
[1094] The server uses an image processing means to read the image file, first convert it to grayscale and remove noise, then process it to distinguish between pictures and text. Next, a character recognition means uses optical character recognition (OCR) technology to extract text from the preprocessed image. This technology can be implemented using common OCR software (e.g., Google OCR).
[1095] The extracted strings are temporarily stored on a server and then translated into the target language by a language translation tool using a generative AI model (e.g., OpenAI's GPT model) to provide highly accurate translation based on context.
[1096] Furthermore, the emotion recognition means analyzes the user's emotion data and adjusts the translation result based on the emotion. For example, if the user expresses a positive emotion, the translation result is adjusted to a more emotional expression.
[1097] Finally, the translation result is passed to the translation result output means, where it is converted into a data format suitable for the user interface. The server sends this response to the terminal, and the terminal displays the translation result to the user through the user interface.
[1098] As a concrete example, if a user translates a Japanese manga page "manga_page.jpg" into the target language, English, and simultaneously inputs a smiley face as emotion data, the process will be as follows: When the user inputs the image file path and the target language "en," the device sends this information and emotion data to the server. The server processes the image, extracts the string "hello" and translates it to "Hello." Because the emotion recognition means identifies the user's positive emotion, the translation result is adjusted to "Hello! How are you feeling today?" This result is sent to the device and displayed to the user.
[1099] Examples of prompts include:
[1100] Analyze the manga image file "manga_page.jpg". Then, translate the Japanese text extracted from this image into English, providing a translation that reflects a positive emotion because the user is smiling. The output should be in the following format: "Hello! How are you feeling today?".
[1101] In this way, the system of the present invention coordinates advanced processes, quickly provides high-quality translation results, and realizes natural, familiar translation that responds to the user's emotions.
[1102] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1103] Step 1:
[1104] The user inputs an image file and a target language into the terminal. Specifically, the user selects an image file called "manga_page.jpg" and English (en) as the target language. The user then inputs their own emotion (e.g., a smile) using a camera or microphone. The input data includes the image file path, the target language, and emotion data.
[1105] Input: Image file "manga_page.jpg", target language "en", emotion data (e.g., smile)
[1106] Output: Request data sent by the device to the server
[1107] Step 2:
[1108] The device sends the input data to the server. Specifically, the device combines the image file path "manga_page.jpg," the target language "en," and the emotion data into a single request data and sends it to the server.
[1109] Input: Request data (image file path, target language, emotion data)
[1110] Output: The request data received by the server
[1111] Step 3:
[1112] The server parses the request data. Specifically, the server receives the request data and separates and obtains the image file path, target language, and emotion data. Parsing is done using methods such as JSON parsing.
[1113] Input: Request data
[1114] Output: Individual data after analysis (image file path, target language, emotion data)
[1115] Step 4:
[1116] The server performs image processing. Specifically, the server reads the image file using image processing means, converts it to grayscale, then removes noise, and then performs processing to distinguish between pictures and text.
[1117] Input: Image file "manga_page.jpg"
[1118] Output: Image data after grayscale conversion and noise removal
[1119] Step 5:
[1120] The server extracts characters using optical character recognition (OCR) technology. Specifically, it extracts characters from the preprocessed image data using OCR technology. This process extracts the string "hello."
[1121] Input: Image data after grayscale conversion and noise removal
[1122] Output: The extracted string "Hello"
[1123] Step 6:
[1124] The server translates the string using a language translation tool. Specifically, it translates the extracted string "hello" into English using a generative AI model (e.g., GPT model). This translation results in the string "Hello."
[1125] Input: Extracted string "Hello"
[1126] Output: The translated string "Hello"
[1127] Step 7:
[1128] The server uses emotion recognition to adjust the translation result. Specifically, it adjusts the translation result based on emotion data (information that the user is smiling). As a result, the translated string "Hello" is changed to "Hello! How are you feeling today?"
[1129] Input: translated string "Hello", emotion data (e.g., smile)
[1130] Output: Adjusted translation result "Hello! How are you feeling today?"
[1131] Step 8:
[1132] The server converts the translation results into an output format. Specifically, it converts the adjusted translation results into a data format (e.g., HTML format) for display in a user interface.
[1133] Input: Adjusted translation result "Hello! How are you feeling today?"
[1134] Output: Data format for user interface display
[1135] Step 9:
[1136] The server generates response data and transmits it to the terminal. Specifically, the server generates response data including the converted translation result and transmits it to the terminal.
[1137] Input: Data format for user interface display
[1138] Output: Response data sent to the device
[1139] Step 10:
[1140] The device receives the response data from the server and displays the translation result to the user. Specifically, the device analyzes the response data received from the server and displays the translation result "Hello! How are you feeling today?" to the user through the user interface.
[1141] Input: Response data from the server
[1142] Output: The translation displayed to the user: "Hello! How are you feeling today?"
[1143] (Application example 2)
[1144] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1145] The present invention aims to improve the quality of communication in a multilingual translation system by not only reducing language barriers between employees but also by providing translation results that take into consideration the feelings of each employee. Another aim is to improve work efficiency by providing appropriate feedback when employees are tired or stressed.
[1146] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1147] In this invention, the server includes image processing means, character recognition means, language translation means, emotion recognition means, and emotion-based translation result adjustment means, which allows for extracting characters from images, translating them into a target language, and detecting the employee's emotional state and adjusting the translation result accordingly.
[1148] "Image processing means" refers to devices or software that have the function of extracting specific information from image files or analyzing images.
[1149] "Character recognition means" means a device or software that has the capability to use optical character recognition technology to detect characters in an image file and convert them into digital text.
[1150] "Language translation means" refers to a device or software capable of translating extracted digital text into another language.
[1151] "Translation result output means" refers to a device or software that has the function of providing an interface for displaying translated text to a user.
[1152] "Emotion recognition means" refers to a device or software that has the function of analyzing and recognizing emotions from a user's facial image or voice data.
[1153] "Means for adjusting translation results based on emotions" refers to devices or software that have the function of adjusting translation results based on recognized emotional data to provide more appropriate and natural expressions.
[1154] The present invention provides a system for supporting multilingual communication among employees in a factory and providing appropriate feedback based on the emotions of the employees. The system includes image processing means, character recognition means, language translation means, translation result output means, emotion recognition means, and emotion-based translation result adjustment means.
[1155] First, the user (employee) inputs the image file of the work instruction (e.g., "work_instruction.jpg") into the terminal. The emotion recognition means analyzes the user's facial image (e.g., "employee_face.jpg") to recognize the emotion. This data is sent to the server via the terminal.
[1156] The server first reads the image file using the image processing means, converts it to grayscale, and removes noise. Next, the character recognition means performs optical character recognition (OCR) to extract characters from the image. This extracted string is temporarily saved and translated into the target language (e.g., English "en") by the language translation means.
[1157] The emotion recognition means analyzes the user's facial image and obtains emotion data (e.g., tired, positive, negative). The emotion-based translation result adjustment means uses this emotion data to adjust the translation result to include appropriate feedback. For example, if the user is recognized as tired, the translation result is adjusted to say something like, "Next, please press this button. Is everything okay?"
[1158] Finally, the adjusted translation result is converted into a display format by the translation result output means, returned to the terminal, and displayed to the user.
[1159] As a concrete example, when a user translates a Japanese work instruction document "work_instruction.jpg" into English and the emotion engine simultaneously recognizes the user's tired expression, the process is as follows: When the user inputs the image file path and the target language "en", the device sends this information and emotion data to the server. The server processes the image, extracts the string "For the next step, please press this button," and translates it as "Next, please press this button." Because the emotion engine recognizes the user's tired expression, the translation result is adjusted to "Next, please press this button. Is everything okay?" This result is sent to the device and displayed to the user.
[1160] An example of an input prompt for a generative AI model is, "Extract Japanese text from the image file 'work_instruction.jpg' and translate it into English. In doing so, output a translation result that includes appropriate feedback based on the employee's emotions read from 'employee_face.jpg'."
[1161] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1162] Step 1:
[1163] The user inputs the image file of the work instruction (e.g., "work_instruction.jpg") and the user's face image (e.g., "employee_face.jpg") into the terminal. At this time, the target language is also input. The input data includes the image file path, the face image file path, and the target language information.
[1164] Step 2:
[1165] The terminal sends the input data (image file path, face image file path, target language) to the server. The server receives this data and passes it to the image processing means and emotion recognition means.
[1166] Step 3:
[1167] The server uses image processing means to read the image file ("work_instruction.jpg"), convert it to grayscale, and remove noise. The input data is the image file, and the output data is the preprocessed image file.
[1168] Step 4:
[1169] The server uses character recognition means to extract characters from the preprocessed image file. Specifically, it converts the characters into digital text using optical character recognition (OCR) technology. The input data is the preprocessed image file, and the output data is the extracted character string.
[1170] Step 5:
[1171] The server uses a language translation mechanism to translate the extracted string into the target language (e.g., English "en"), using generative artificial intelligence to provide a context-appropriate translation. The input data is the extracted string, and the output data is the translated string.
[1172] Step 6:
[1173] The server uses emotion recognition means to analyze the user's facial image ("employee_face.jpg") and recognize the user's emotional state. The input data is a facial image file, and the output data is recognized emotional data.
[1174] Step 7:
[1175] The server adjusts the translation result based on the recognized emotion data using a means for adjusting the translation result based on emotion. For example, if the server recognizes that the user is tired, it adds thoughtful feedback to the translation result. The input data is the translated string and emotion data, and the output data is the adjusted translation result.
[1176] Step 8:
[1177] The adjusted translation result is converted into a display format by the translation result output means and sent to the terminal. The terminal receives this data and displays it to the user. The input data is the adjusted translation result, and the output data is in a display format that the user can see.
[1178] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1179] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1180] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1181] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1182] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1183] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1184] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1185] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, motorcycles, and other devices, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1186] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1187] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1188] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1189] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1190] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1191] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1192] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1193] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1194] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1195] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1196] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1197] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1198] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1199] The following is further disclosed regarding the above embodiment.
[1200] (Claim 1)
[1201] image processing means;
[1202] character recognition means;
[1203] a language translation means;
[1204] A translation result output means;
[1205] Multilingual translation system including.
[1206] (Claim 2)
[1207] 2. The multilingual translation system according to claim 1, wherein the image processing means distinguishes between pictures and text from the image file.
[1208] (Claim 3)
[1209] 2. The multilingual translation system according to claim 1, wherein the character recognition means extracts characters using optical character recognition technology.
[1210] (Claim 4)
[1211] 2. The multilingual translation system according to claim 1, wherein the language translation means performs translation using generative artificial intelligence.
[1212] (Claim 5)
[1213] 2. The multilingual translation system according to claim 1, wherein the translation result output means displays the translation result on a user interface.
[1214] (Claim 6)
[1215] 2. The multilingual translation system according to claim 1, wherein the image processing means is adapted to a particular genre or style of manga.
[1216] "Example 1"
[1217] (Claim 1)
[1218] image processing means;
[1219] character recognition means;
[1220] a language translation means;
[1221] A translation result output means;
[1222] a terminal on which a user inputs an image file and a target language;
[1223] A server that receives and processes information from the terminal;
[1224] A system including:
[1225] (Claim 2)
[1226] 2. The system according to claim 1, wherein the image processing means distinguishes between pictures and text from the image file, and performs grayscale conversion and noise removal.
[1227] (Claim 3)
[1228] 2. The system of claim 1, wherein the character recognition means extracts characters using optical character recognition techniques.
[1229] (Claim 4)
[1230] 2. The system of claim 1, wherein the language translation means uses generative artificial intelligence to translate the extracted text into the target language.
[1231] "Application Example 1"
[1232] (Claim 1)
[1233] image processing means;
[1234] character recognition means;
[1235] a language translation means;
[1236] A translation result output means;
[1237] Smart glasses and
[1238] A means for translating images captured by smart glasses in real time;
[1239] A system including:
[1240] (Claim 2)
[1241] 2. The system according to claim 1, wherein the image processing means distinguishes between pictures and text from the image file.
[1242] (Claim 3)
[1243] 2. The system of claim 1, wherein the character recognition means extracts characters using optical character recognition techniques.
[1244] "Example 2: Combining Emotion Engines"
[1245] (Claim 1)
[1246] image processing means;
[1247] character recognition means;
[1248] a language translation means;
[1249] A translation result output means;
[1250] An emotion recognition means;
[1251] user interface means;
[1252] A system including:
[1253] (Claim 2)
[1254] 2. The system according to claim 1, wherein the image processing means distinguishes between pictures and text from the image file.
[1255] (Claim 3)
[1256] 2. The system of claim 1, wherein the character recognition means extracts characters using optical character recognition techniques.
[1257] (Claim 4)
[1258] 2. The system according to claim 1, wherein the language translation means performs translation using a generative AI model.
[1259] (Claim 5)
[1260] 2. The system according to claim 1, wherein the emotion recognition means analyzes the emotion of the user and adjusts the translation result based on the emotion.
[1261] (Claim 6)
[1262] 2. The system according to claim 1, wherein the translation result output means converts the translation result into a data format suitable for display on a user interface.
[1263] "Application example 2 when combining emotion engines"
[1264] (Claim 1)
[1265] image processing means;
[1266] character recognition means;
[1267] a language translation means;
[1268] A translation result output means;
[1269] An emotion recognition means;
[1270] a means for adjusting the translation result based on emotion;
[1271] A system including:
[1272] (Claim 2)
[1273] 2. The system according to claim 1, wherein the image processing means distinguishes between pictures and text from the image file.
[1274] (Claim 3)
[1275] 2. The system of claim 1, wherein the character recognition means extracts characters using optical character recognition techniques. [Explanation of symbols]
[1276] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. image processing means; character recognition means; a language translation means; A translation result output means; Multilingual translation system including.
2. 2. The multilingual translation system according to claim 1, wherein the image processing means distinguishes between pictures and characters from the image file.
3. 2. The multilingual translation system according to claim 1, wherein the character recognition means extracts characters using optical character recognition technology.
4. 2. The multilingual translation system according to claim 1, wherein the language translation means performs translation using generative artificial intelligence.
5. 2. The multilingual translation system according to claim 1, wherein the translation result output means displays the translation result on a user interface.
6. 2. The multilingual translation system according to claim 1, wherein the image processing means is adapted to a particular genre or style of manga.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A