system
The system addresses the language barrier in educational videos by translating audio and text in real time, ensuring users can access high-quality content in their native language, thus promoting multicultural education.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
The language barrier in educational videos limits global access to high-quality learning content, and existing solutions fail to provide real-time multilingual translation for both audio and visual text information.
A system that translates audio and text information from educational videos in real time using speech recognition, optical character recognition, and multilingual translation technologies, synchronizing the translated content with the original video and delivering it to users through streaming.
Enables users to access educational content in their native language, overcoming language barriers and providing a multicultural learning environment.
Smart Images

Figure 2026070182000001_ABST
Abstract
Description
Technical Field
[0001] The technology of this disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Currently, the language barrier in videos of lectures and classes has become an obstacle to global education access. In particular, for users who speak different languages, there is a problem that the methods for accessing high-quality learning content translated in real time are limited. In addition, existing solutions that provide real-time multilingual translation not only for the audio in the video but also for visual text information are insufficient. Therefore, there is a need for a system to facilitate international access to education and information and enable users of all languages to obtain a similar learning experience.
Means for Solving the Problems
[0005] This invention provides a means for translating audio data contained in online lectures and class videos from a first language to a second language in real time. Furthermore, it uses optical character recognition technology to extract and translate text information displayed within the video data. These translation results are synchronized with the original video to construct a generated video containing the translated audio and text. Since the generated video is streamed to users in real time, it provides an environment in which viewers using various languages can enjoy equivalent educational content, overcoming language barriers.
[0006] "Real-time" refers to a process or operation that takes place almost simultaneously with data input, with the results being output immediately.
[0007] "Audio data" refers to information that digitally represents the speaker's voice or sounds, collected by devices such as microphones.
[0008] The "primary language" refers to the source language, which is the language of the original audio or text being processed.
[0009] The "second language" refers to the target language of the translation, the language in which the converted audio or text is expressed.
[0010] "Translation" refers to the process of converting audio or text expressed in one language into a different language.
[0011] "Video data" refers to dynamic visual information captured by cameras, recording devices, etc., expressed in digital format.
[0012] "Text information" refers to data expressed in the form of sentences or characters, and includes text information displayed in documents, on screens, etc.
[0013] Optical character recognition (OCR) technology is a technology that automatically detects characters contained in images and videos and converts them into digital text.
[0014] "Generated video" refers to video that has been processed from the original video to include new information (such as translated audio or text).
[0015] "Streaming distribution" refers to a method of providing digital media by transferring it in real time over the internet, allowing viewers to watch it simultaneously without waiting for downloads. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention provides a system that translates lectures and class videos into multiple languages in real time and delivers them to users. This system operates through a server, user terminals, and a digital network.
[0038] The user's device functions as a device that sends viewing requests to the server. The user selects the lecture or lesson video they want to watch and sends the request to the server. At this stage, they can specify their desired translation language.
[0039] The server receives requests from users and retrieves videos of selected lectures or classes via the network. Using speech recognition technology, the server analyzes the audio track of the video and converts it into text. This transcribed audio data is then translated in real time by the server.
[0040] Simultaneously, the server uses optical character recognition technology to extract text written on whiteboards or blackboards in the video. This visual text information is also translated into the selected language.
[0041] The translated audio and text are reconstructed on the server and overlaid onto the original video. For example, if a math lesson conducted in English is translated into Japanese, the English lecture audio is replaced with Japanese audio, and the mathematical formulas and explanations written on the whiteboard are translated into Japanese and displayed on the video. This process allows users to view learning content translated into their native language in real time.
[0042] The user's device functions as a device that receives translated video streamed from the server and plays it back in real time. On the device, the video and audio are synchronized and adjusted to flow without delay.
[0043] In this way, the present invention provides a system that transcends language barriers, offers a multicultural learning environment, and enables users worldwide to access high-quality educational content.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] Users send a request from their device to the server to view a video of a specific lecture or class. They also specify their preferred translation language at the same time.
[0047] Step 2:
[0048] The server receives a request from the user and connects to the specified video stream. It then prepares to retrieve the streamed video.
[0049] Step 3:
[0050] The server analyzes the audio track of the acquired video using speech recognition technology and converts the audio data into text. This transcribed audio data is then used in the translation process.
[0051] Step 4:
[0052] The server analyzes the video frames and uses optical character recognition technology to extract text information written on the whiteboard or blackboard. This information is also processed as text data.
[0053] Step 5:
[0054] The server translates both audio-text and video-text into a specified second language. This translation utilizes generative AI technology to achieve real-time processing.
[0055] Step 6:
[0056] The server uses speech synthesis technology to convert the translated audio text into speech in the specified language. The translated text is also synchronized with the original video, overlaid on top of it.
[0057] Step 7:
[0058] The server generates a video that integrates the translated audio and text, and streams it to the user's device in real time.
[0059] Step 8:
[0060] The user's device receives translated video transmitted from the server and plays it back in real time. Audio and video are synchronized, providing the user with a lag-free viewing experience.
[0061] (Example 1)
[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0063] In modern society, multilingual translation is crucial for people who speak different languages to smoothly share and understand information. However, technologies for translating lectures and class videos in real time face many challenges, including accurate synchronization of audio and text, translation accuracy and speed, and smooth delivery to users. This invention aims to solve these problems and provide a system that promotes multicultural education and knowledge sharing by removing barriers between different languages.
[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0065] In this invention, the server includes means for receiving requests to view lectures or classes from user terminals, means for acquiring video via the network based on the request and converting the audio track into text using speech recognition technology, and means for extracting text information from the video using optical character recognition technology. This enables the content of lectures and classes provided in different languages to be translated quickly and accurately into other languages, allowing users to view videos in their own language in real time.
[0066] A "user terminal" is a device used by a user to retrieve information from a server via the internet and to view video and audio.
[0067] A "server" is a computer system that receives requests from users on a network, processes the data, and provides the results to the user.
[0068] A "viewing request" is a request that a user sends to a server to view a video of a specific lecture or class.
[0069] A "network" is a system that connects computers and devices to each other, enabling the transmission and reception of data.
[0070] "Speech recognition technology" is a general term for the processes and technologies used to convert speech data into text data.
[0071] Optical Character Recognition (OCR) is a technology that extracts character information from image data, and is often abbreviated as OCR.
[0072] "Multilingual translation technology" is a system of technologies for converting text or audio in one language into another language.
[0073] "Speech synthesis technology" is a technology that converts text into audio data and reproduces it audibly.
[0074] "Means of generating video" refers to the process of integrating translated audio and text to create new video data.
[0075] "Means of real-time delivery" refers to technologies and methods that enable the immediate transmission of generated video to the user's device, allowing for viewing without delay.
[0076] This invention is a system that translates lectures and class videos into multiple languages in real time and provides them to users. The system consists of a server, user terminals, and a digital network connecting them.
[0077] The server is a computer system that receives viewing requests from users, which include a video ID and the desired translation language. Upon recognizing the request, the server retrieves the relevant video metadata from the database and pulls the video file from the appropriate storage system. As speech recognition technology, the Google® Speech-to-Text API can be used to convert the audio track into text data.
[0078] Furthermore, Tesseract OCR is applied to optical character recognition technology to extract text information from the video. This text is then translated in real time into the specified language using multilingual translation technologies such as the Google Cloud Translation API. The translated text is then generated as an audio file using speech synthesis technology (e.g., Google Text-to-Speech API).
[0079] The user terminal receives translated video streamed from the server and provides the video to the user through sight and sound. The terminal functions as a media player, decoding and buffering the received data and adjusting it so that the video and audio are played in sync in real time.
[0080] For example, if a history lesson conducted in English is translated into French, and the user requests French on their device, the server can retrieve the English lecture video, translate the audio and text data into French in real time, and provide it to the user as a streaming video. An example of a prompt related to this process would be, "Please provide a real-time translation of an English history lecture into French, including both the spoken lecture and any written content on blackboards, making it viewable in a video format."
[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0082] Step 1:
[0083] The user uses their device to request to view a video of a lecture or class. Specifically, they enter the video title and desired translation language into the user interface and press the "Submit Request" button. At this time, the video ID and translation language data are sent to the server as input. The output is the request data received by the server.
[0084] Step 2:
[0085] The server parses the request received from the user and retrieves metadata from the database based on the video ID. This data includes links to the necessary video files based on the request. The input is the request data, and the output is the retrieved video metadata.
[0086] Step 3:
[0087] The server downloads the relevant video file from the storage system. This video file contains an audio track, which is converted to text by a speech recognition system. Specifically, the Google Speech-to-Text API receives the audio data and generates the corresponding text data. The input is the video file, and the output is the transcribed audio data.
[0088] Step 4:
[0089] The server extracts text information from video frames using optical character recognition (OCR) technology. Video frames are captured and converted into text information by the Tesseract OCR engine. The input is the video frame, and the output is the extracted text data.
[0090] Step 5:
[0091] The server translates the transcribed audio data and extracted text data into the specified language using the Google Cloud Translation API. This involves applying a translation algorithm to the data and converting it to the specified target language. The input is text data, and the output is translated text data.
[0092] Step 6:
[0093] The server synthesizes the translated text data into speech using the Google Text-to-Speech API and generates an audio file. The input is the translated text data, and the output is the synthesized audio file.
[0094] Step 7:
[0095] The server integrates the translated audio and text into the original video data to generate a new video. Video editing software then edits the video to ensure that subtitles and audio are displayed correctly. The input consists of translated audio files and text data, while the output is the generated video data.
[0096] Step 8:
[0097] The server streams the generated video to the user's terminal in real time. The terminal synchronizes the received video data, making it viewable to the user without delay. The input is the generated video data, and the output is the video played on the terminal.
[0098] (Application Example 1)
[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0100] In conventional systems, real-time multilingual translation of lectures and classes was difficult, creating a significant barrier to access for users in different language environments. Furthermore, the lack of visualization of audio and text information within videos limited the user's viewing experience. This hindered the equalization of education in a multicultural society.
[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0102] In this invention, the server includes means for processing audio data in real time and translating it from a first language to a second language; means for extracting text information from video data and translating it from a first language to a second language; means for synchronizing the translation results with the video data and constructing generated video including the translated audio and text; communication means for delivering the generated video to a terminal in real time and enabling the user to specify a selective viewing language; and means for providing educational information in the user's native language in the selected language and displaying it via a visual device. This enables the user to overcome language barriers and view high-quality educational content in their native language in real time.
[0103] "Real-time" means that information processing is performed instantly on the spot.
[0104] "Audio data" refers to audio information represented in digital format.
[0105] The "primary language" is the language in which the original spoken or written text is expressed.
[0106] A "second language" is the language in which the translated information is intended to be expressed.
[0107] "Means of translation" refer to functions or processes for converting information from one language into another language.
[0108] "Video data" refers to digital data that includes visual information.
[0109] "Text information" refers to information expressed in the form of characters.
[0110] "Synchronizing" means making different pieces of information or actions coincide and match at the same time.
[0111] "Generated video" refers to video that has been newly created by processing the original video.
[0112] A "terminal" is an information processing device used by a user.
[0113] "To distribute" means to send information to a recipient.
[0114] "Communication methods" refer to the methods and technologies used to send and receive data.
[0115] "Auditory language" refers to language intended for users to understand through their sight and hearing.
[0116] "Educational information" refers to information that contains knowledge and content useful for learning.
[0117] A "visual device" is a hardware device that displays video information.
[0118] This invention describes a system for delivering educational content translated into multiple languages in real time using specific devices and software.
[0119] The server first converts the audio data into text using speech recognition technology. Cloud-based services such as Google Cloud Speech-to-Text or Amazon Transcribe can be used for speech recognition. This text is then translated into a second language by a translation engine. Multilingual translation services, such as the Google Translate API, are used in the translation process.
[0120] The server also uses optical character recognition (OCR) technology to extract text information from the video data. In this process, the text data on the whiteboard or blackboard is converted into a digital format and translated into the specified language. Using video processing software such as FFmpeg (described later), the translated text and audio are accurately overlaid onto the original video to construct the generated video.
[0121] The generated video is delivered to the user's device in real time. This device may be a smartphone, smart glasses, or a head-mounted display. The delivered video can be viewed in the user's chosen language, and playback is achieved with minimal latency.
[0122] For example, when a user watches a lecture given at a university in a different country, this system translates the lecture audio into the user's native language and also translates and displays the contents of the whiteboard, thereby providing equal educational opportunities that transcend language barriers.
[0123] An example of a prompt for a generative AI model might be: "Use a translation service to convert the lecture audio from English to Japanese in real time. Additionally, translate the text in the video into Japanese and display it in sync with the video."
[0124] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0125] Step 1:
[0126] The user uses their device to select the lecture or lesson content they wish to view and their preferred translation language, then submits a request. The user's selection information is used as input and sent to the server. The output consists of content identification information and the translation language setting.
[0127] Step 2:
[0128] The server retrieves the specified content via the network based on a request received from the user. The input is the user's request information, and the output is the audio and video data of the content. This process prepares the data to be viewed within the system.
[0129] Step 3:
[0130] The server uses speech recognition technology to convert the acquired audio data into text data. The input is audio data, and the output is text data converted from the audio. In this step, a service such as Google Cloud Speech-to-Text is used to convert the audio to text.
[0131] Step 4:
[0132] The server uses optical character recognition (OCR) technology to extract text information from video data. The input is video data, and the output is extracted text data. For example, OCR software such as Tesseract may be used.
[0133] Step 5:
[0134] The server translates text data into a specified second language. Input is text data obtained from audio and video, and output is translated text data. Real-time translation is performed using APIs such as Google Translate.
[0135] Step 6:
[0136] The server constructs a generated video by overlaying the translated audio and text data onto the original video data. The input is the translated audio and text along with the original video data, and the output is a viewable generated video. Video editing software such as FFmpeg is used to integrate the audio and text into the video.
[0137] Step 7:
[0138] The server streams the generated video to the user's terminal in real time. The input is the generated video data, and the output is the video played on the user's terminal. On the terminal side, the translated content is played back without delay.
[0139] Step 8:
[0140] Users can view lectures and lessons in their chosen language on their devices, enabling them to understand content in their native language. Users can also receive educational information in real time through streamed, generated video.
[0141] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0142] This invention provides a system that translates the audio and video content of lectures and classes into multiple languages in real time, and also has the function of recognizing the user's emotions. This system operates via a server, user terminals, an emotion engine, and a digital network connecting them.
[0143] The user's device is the one that sends viewing requests to the server. The user selects the lecture or lesson video they want to watch, specifies their desired translation language, and sends the request to the server.
[0144] The server receives a request from the user, retrieves the video stream over the network, and begins processing. Speech recognition technology is applied on the server, converting the video's audio track into text data. This text data is then translated into the specified language using generative AI.
[0145] For the visual elements of the video, the server uses optical character recognition technology to extract text information and translates it in the same way. The translation results are converted into speech in the specified language using speech synthesis technology and overlaid on the video.
[0146] Furthermore, the emotion engine analyzes the user's emotions in real time using data from the device's built-in camera or data provided by the user. This emotion data is sent to the server and influences the display of translated content. For example, if the emotion engine determines that the user is confused, the server can enlarge the explanatory text or slow down the playback of the audio commentary.
[0147] The user's device receives translated and emotion-adapted video streamed from the server and plays it back in real time. This allows users to view translated content in a way that is easier to understand and responds to their individual emotional state.
[0148] In this way, the present invention overcomes language barriers and improves the user experience, thereby providing a multicultural learning environment and realizing a system that enables users worldwide to access high-quality educational content.
[0149] The following describes the processing flow.
[0150] Step 1:
[0151] The user uses their device to select the lecture or lesson video they want to watch and sends a request to the server. This request includes the language in which they want the video translated.
[0152] Step 2:
[0153] The server receives a request from the user, connects to the video stream of the specified lecture or class, and prepares to retrieve the data.
[0154] Step 3:
[0155] The server acquires audio data from the video and uses speech recognition technology to convert this audio into text. The resulting text is then used in the translation process.
[0156] Step 4:
[0157] The server analyzes the video data and extracts visual text information from the video using optical character recognition technology. This text information is immediately subjected to a translation process.
[0158] Step 5:
[0159] The server translates audio text data and in-video text information into a second language specified by the user. This translation is optimized for real-time operation.
[0160] Step 6:
[0161] The server converts the translated text information into speech using speech synthesis technology and places it as subtitles on the video. This synthesized speech and text are synchronized with the original video.
[0162] Step 7:
[0163] The user's device captures their facial expressions with a camera, and an emotion engine analyzes the user's emotions in real time based on this data. The emotion engine then sends the analysis results to the server.
[0164] Step 8:
[0165] Based on data received from the emotion engine, the server adaptively adjusts the translated content according to the user's emotions. For example, if the user is confused, the server will deliver content that emphasizes explanatory parts or plays audio at a slower pace.
[0166] Step 9:
[0167] The server streams the adjusted and translated video to the user's device in real time.
[0168] Step 10:
[0169] The user's device receives translated video streamed from the server and plays it back while synchronizing the video and audio. This allows the user to instantly view content adjusted to their own language.
[0170] (Example 2)
[0171] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0172] It is necessary to remove language barriers in international lectures and classes and provide support to help participants gain a better understanding. Furthermore, there is a need to improve learning effectiveness by providing flexible content that responds to participants' emotional states.
[0173] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0174] In this invention, the server includes means for converting audio data into text data in real time using speech recognition technology and translating the text data from a first language to a second language; means for extracting character information from video data using optical character recognition technology and translating the extracted character information from a first language to a second language; means for converting the translation result into audio data using speech synthesis technology and constructing a generated video by overlaying the audio data and translated text onto video data; means for adjusting the display timing and method of the generated video based on emotion data obtained from a device that analyzes the user's emotions; and means for delivering the generated video to an information terminal in real time. As a result, the user can receive translated content adapted to their emotional state in real time.
[0175] "Speech recognition technology" is a technology that converts human speech into text data using a computer system.
[0176] "Optical character recognition technology" is a technology that extracts character information from image or video data.
[0177] "Translation means" refers to a process or device for converting text data into another language.
[0178] "Speech synthesis technology" is a technology that converts text data into speech data and outputs it as natural-sounding speech.
[0179] A "generative AI model" is an artificial intelligence system that performs natural language processing to generate and translate text data.
[0180] A "sentiment analysis device" is a device that analyzes a user's emotions in real time and acquires that data.
[0181] "Video data" refers to digital data that contains visual information.
[0182] An "information terminal" is an electronic device that transmits, receives, and processes data via a network.
[0183] This invention relates to a system that translates the audio and video content of lectures and classes into other languages in real time and also has the function of recognizing the user's emotions. This system consists of a server, information terminals, an emotion analysis device, and a network connecting them.
[0184] The user first uses an information terminal to select the content they want to view and specify the desired translation language. This request is sent to the server. The server uses speech recognition technology (e.g., commercial speech recognition software) to convert the acquired speech data into text data.
[0185] Next, a generative AI model (for example, a general-purpose generative AI engine) is used to translate the text data into the specified language. A prompt such as "Translate Japanese into English" is used.
[0186] For video data, the server extracts text information using optical character recognition technology (e.g., common OCR software) and translates this information. The translation result is converted into speech using speech synthesis technology (e.g., standard text-to-speech software) and incorporated into the video as an overlay.
[0187] For sentiment analysis, the camera and microphone built into the user's information terminal are used. The data acquired is analyzed in real time by a sentiment analysis device (e.g., a standard sentiment recognition API), and the sentiment information is sent to a server. Based on the sentiment information, the server can dynamically adjust how the generated video is displayed.
[0188] A concrete example is a user who wants to understand a Spanish physics lecture in English. The user selects the relevant Spanish lecture from their device and specifies English as the translation language. The audio and video are recognized and translated, and the audio and text are provided to the user in English in a synchronized format. Furthermore, the server can appropriately adjust the playback speed and level of detail according to the user's level of difficulty.
[0189] In this way, this system enables users to understand educational content in other languages in a way that is optimized for their own emotions.
[0190] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0191] Step 1:
[0192] The user selects the lecture or class content they wish to view using their information terminal and specifies their desired translation language. A request is generated containing the user's selected video data and desired language information. This request is sent to the server, where processing begins.
[0193] Step 2:
[0194] The server retrieves a specified video stream over the network. It uses video data based on user requests as input. The output is generated by converting audio data into text data using speech recognition technology. Specifically, the speech recognition engine analyzes the audio waveform and outputs it as text in the corresponding language.
[0195] Step 3:
[0196] The server uses a generative AI model to translate the text data obtained in step 2 into the specified language. The recognized original text data and a prompt message are used as input. A prompt message such as "Please translate Japanese into English" is used. The output is the translated text data.
[0197] Step 4:
[0198] The server extracts text information from video data using optical character recognition (OCR) technology. The input is video data, and the output is the extracted text information. The OCR engine analyzes the characters in the image and outputs them as corresponding text.
[0199] Step 5:
[0200] The character information extracted in Step 4 is similarly translated into the specified language using the generative AI model. The recognized character information and prompt text are used as input, and the output is the translated character information.
[0201] Step 6:
[0202] The server uses speech synthesis technology to convert translated text data into new speech data. The input is translated text data, and the output is speech data in the specified language. The speech synthesis engine analyzes the text and outputs it as synthesized speech.
[0203] Step 7:
[0204] The server overlays the generated audio data and translated text onto the video data to construct the generated video. The input consists of the generated audio data, translated text data, and original video data, and the output is the generated video in which they are integrated.
[0205] Step 8:
[0206] The emotion analysis device uses data obtained from the camera and microphone built into the device to detect the user's emotions and sends it to the server as temporary input data. This data indicates the emotional state, and a specific emotional pattern is obtained as output.
[0207] Step 9:
[0208] The server adjusts how the generated video is displayed based on the user's emotion data. The input is emotion data and generated video data, and the output is the generated video with the adjusted playback parameters.
[0209] Step 10:
[0210] The server delivers the adjusted, generated video to the information terminal in real time. Users view the content played on their terminal, receiving translations and emotionally-responsive learning support. The final output is the content played back in real time.
[0211] (Application Example 2)
[0212] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0213] In an increasingly globalized world, understanding educational and informational content provided in different languages remains challenging, transcending language barriers. Furthermore, the lack of content adjustments tailored to the audience's emotions and level of understanding leads to decreased learning efficiency. Therefore, a system is needed that provides real-time multilingual translation and customized information based on the audience's emotions.
[0214] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0215] In this invention, the server includes means for processing audio data in real time and translating it from a first language to a second language, means for extracting text information from video data and translating it from a first language to a second language, and means for acquiring the user's emotional state and adjusting the display method of the generated video based on the emotion. This enables the display of multilingual educational content in a customized manner according to the viewer's emotions, resulting in a more effective learning experience.
[0216] "Means for processing audio data in real time and translating from a first language to a second language" refers to a device or function that performs the process of instantly analyzing an audio signal, converting it into text, and then translating that text into a specified different language.
[0217] "Means for extracting text information from video data and translating it from a first language to a second language" refers to a device or function that identifies and acquires characters within a video and processes those characters to represent them in a different language.
[0218] "Means for synchronizing translation results with video data and constructing generated video including translated audio and text" refers to a device or function that performs the process of combining audio and text translated into different languages with the original video to create integrated video content.
[0219] "Means for delivering generated video to terminals in real time" refers to a network system and its control process that transfers the created translated video to the viewer's device without delay.
[0220] "Means for acquiring the user's emotional state and adjusting the display method of generated video based on that emotion" refers to a device or function that analyzes the user's facial expressions and behavioral data to recognize emotions and automatically changes the video playback speed, display content, etc., according to the analysis results.
[0221] The system for realizing this invention is composed of an integrated set of technologies, primarily consisting of servers, user terminals, and various software technologies.
[0222] First, the server receives a viewing request from the user. This request includes the content to be viewed and the desired translation language. The server acquires the audio data and converts it to text using speech recognition technology. This process primarily uses Python and speech recognition libraries.
[0223] Furthermore, the server uses optical character recognition (OCR) technology to extract text information from the video data. Simultaneously, a generative AI model is used to translate this text into the language specified by the user. At this stage, the OpenAI® API is often used.
[0224] The server uses the user's device camera to analyze facial expressions and behavior in real time in order to acquire the user's emotional state. Machine learning libraries such as TENSORFLOW® are used for emotion analysis.
[0225] The converted audio and text are synchronized with the video data as translation results. Speech synthesis technology is used here, and the resulting video is displayed in a way that adjusts based on the user's emotional state. For example, if the user is determined to be confused, the playback speed can be adjusted, or additional detailed explanations can be displayed. Finally, the adjusted video and audio are delivered to the device.
[0226] For example, if a Japanese user watching a technical lecture in English uses this system, and it determines that technical terms are difficult to understand, the system will display a detailed explanation of those terms in Japanese and play the audio commentary at a slower pace. The following is an example of a prompt to the generative AI model:
[0227] "Translate the following educational content from English to Japanese, emphasizing clarity for a high school student: English lesson content"
[0228] This makes it possible to provide a high-quality learning experience that transcends language and cultural barriers.
[0229] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0230] Step 1:
[0231] The server receives viewing requests from users. The input includes the ID of the content they want to view and the desired translation language. Based on this information, the server retrieves the corresponding video and audio data.
[0232] Step 2:
[0233] The server sends the acquired audio data to the speech recognition engine, which converts it into text data. The input is an audio signal, and the output is the recognized text. This process is performed by a Python speech recognition library.
[0234] Step 3:
[0235] The server extracts text information from video data using optical character recognition (OCR) technology. The input is video frames, and the output is the string data contained within the video. An OCR library is used in this step.
[0236] Step 4:
[0237] The server sends text data to a generation AI model for translation into the desired language. The input is text data in the original language, and the output is text data translated into the specified language. This translation process uses the OpenAI API.
[0238] Step 5:
[0239] The user's device captures facial expressions using its built-in camera and sends them to an emotion recognition engine. The input is facial image data, and the output is data indicating the user's emotional state. Libraries such as TensorFlow are used for emotion analysis.
[0240] Step 6:
[0241] The server synchronizes the translation results with the video data based on sentiment data and adjusts the display accordingly. The input consists of the translation results and sentiment data, and the output is a video with adjusted display settings. Specifically, this includes changes to font size and playback speed.
[0242] Step 7:
[0243] The server delivers generated, translated, and emotion-sensitive video to the terminal. The input is the adjusted video, and the output is this video displayed in real time on the user's terminal. This utilizes streaming technology over a network.
[0244] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0245] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0246] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0247] [Second Embodiment]
[0248] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0249] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0250] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0251] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0252] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0253] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0254] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0255] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0256] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0257] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0258] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0259] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0260] This invention provides a system that translates lectures and class videos into multiple languages in real time and delivers them to users. This system operates through a server, user terminals, and a digital network.
[0261] The user's device functions as a device that sends viewing requests to the server. The user selects the lecture or lesson video they want to watch and sends the request to the server. At this stage, they can specify their desired translation language.
[0262] The server receives requests from users and retrieves videos of selected lectures or classes via the network. Using speech recognition technology, the server analyzes the audio track of the video and converts it into text. This transcribed audio data is then translated in real time by the server.
[0263] Simultaneously, the server uses optical character recognition technology to extract text written on whiteboards or blackboards in the video. This visual text information is also translated into the selected language.
[0264] The translated audio and text are reconstructed on the server and overlaid onto the original video. For example, if a math lesson conducted in English is translated into Japanese, the English lecture audio is replaced with Japanese audio, and the mathematical formulas and explanations written on the whiteboard are translated into Japanese and displayed on the video. This process allows users to view learning content translated into their native language in real time.
[0265] The user's device functions as a device that receives translated video streamed from the server and plays it back in real time. On the device, the video and audio are synchronized and adjusted to flow without delay.
[0266] In this way, the present invention provides a system that transcends language barriers, offers a multicultural learning environment, and enables users worldwide to access high-quality educational content.
[0267] The following describes the processing flow.
[0268] Step 1:
[0269] Users send a request from their device to the server to view a video of a specific lecture or class. They also specify their preferred translation language at the same time.
[0270] Step 2:
[0271] The server receives a request from the user and connects to the specified video stream. It then prepares to retrieve the streamed video.
[0272] Step 3:
[0273] The server analyzes the audio track of the acquired video using speech recognition technology and converts the audio data into text. This transcribed audio data is then used in the translation process.
[0274] Step 4:
[0275] The server analyzes the video frames and uses optical character recognition technology to extract text information written on the whiteboard or blackboard. This information is also processed as text data.
[0276] Step 5:
[0277] The server translates both audio-text and video-text into a specified second language. This translation utilizes generative AI technology to achieve real-time processing.
[0278] Step 6:
[0279] The server uses speech synthesis technology to convert the translated audio text into speech in the specified language. The translated text is also synchronized with the original video, overlaid on top of it.
[0280] Step 7:
[0281] The server generates a video that integrates the translated audio and text, and streams it to the user's device in real time.
[0282] Step 8:
[0283] The user's device receives translated video transmitted from the server and plays it back in real time. Audio and video are synchronized, providing the user with a lag-free viewing experience.
[0284] (Example 1)
[0285] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".
[0286] In modern society, multilingual translation is important for people speaking different languages to smoothly share and understand information. However, technologies for real-time translation of lecture and class videos have many problems such as accurate synchronization of audio and text, translation accuracy and speed, and smooth delivery to users. The object of this invention is to provide a system that solves such problems and promotes multicultural education and knowledge sharing by removing the barriers between different languages.
[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0288] In this invention, the server includes means for receiving a viewing request for a lecture or class from a user terminal, means for acquiring a video via a network based on the request and converting an audio track into text using speech recognition technology, and means for extracting character information from the video using optical character recognition technology. As a result, the content of lectures and classes provided in different languages can be quickly and accurately translated into other languages, and users can view videos that can be understood in their native language in real time.
[0289] The "user terminal" is a device for a user to obtain information from a server via the Internet and view videos and audio.
[0290] The "server" is a computer system for receiving requests from users on a network, performing data processing, and providing the results to users.
[0291] A "viewing request" is a request that a user sends to a server to view a video of a specific lecture or class.
[0292] A "network" is a system that connects computers and devices to each other, enabling the transmission and reception of data.
[0293] "Speech recognition technology" is a general term for the processes and technologies used to convert speech data into text data.
[0294] Optical Character Recognition (OCR) is a technology that extracts character information from image data, and is often abbreviated as OCR.
[0295] "Multilingual translation technology" is a system of technologies for converting text or audio in one language into another language.
[0296] "Speech synthesis technology" is a technology that converts text into audio data and reproduces it audibly.
[0297] "Means of generating video" refers to the process of integrating translated audio and text to create new video data.
[0298] "Means of real-time delivery" refers to technologies and methods that enable the immediate transmission of generated video to the user's device, allowing for viewing without delay.
[0299] This invention is a system that translates lectures and class videos into multiple languages in real time and provides them to users. The system consists of a server, user terminals, and a digital network connecting them.
[0300] The server is a computer system that receives viewing requests from users, which include a video ID and the desired translation language. Upon recognizing the request, the server retrieves the relevant video metadata from the database and pulls the video file from the appropriate storage system. For speech recognition technology, the Google Speech-to-Text API can be used to convert the audio track into text data.
[0301] Furthermore, Tesseract OCR is applied to optical character recognition technology to extract text information from the video. This text is then translated in real time into the specified language using multilingual translation technologies such as the Google Cloud Translation API. The translated text is then generated as an audio file using speech synthesis technology (e.g., Google Text-to-Speech API).
[0302] The user terminal receives translated video streamed from the server and provides the video to the user through sight and sound. The terminal functions as a media player, decoding and buffering the received data and adjusting it so that the video and audio are played in sync in real time.
[0303] For example, if a history lesson conducted in English is translated into French, and the user requests French on their device, the server can retrieve the English lecture video, translate the audio and text data into French in real time, and provide it to the user as a streaming video. An example of a prompt related to this process would be, "Please provide a real-time translation of an English history lecture into French, including both the spoken lecture and any written content on blackboards, making it viewable in a video format."
[0304] The flow of the specific process in Example 1 will be described with reference to FIG. 11.
[0305] Step 1:
[0306] The user uses the terminal to request to view a video of a lecture or a class. Specifically, the user inputs the video title and the desired translation language into the user interface and presses the "Request Submission" button. At this time, the video ID and the translation language data are sent to the server as input. The output is the request data received by the server.
[0307] Step 2:
[0308] The server analyzes the request received from the user and obtains metadata from the database based on the video ID. This data includes a link to the necessary video file based on the request. The input is the request data, and the output is the obtained video metadata.
[0309] Step 3:
[0310] The server downloads the corresponding video file from the storage system. This video file includes an audio track, which is converted into text by an automatic speech recognition system. Specifically, the Google Speech-to-Text API receives the audio data and generates the corresponding text data. The input is the video file, and the output is the texturized audio data.
[0311] Step 4:
[0312] The server extracts character information from the video frames using optical character recognition technology. The video frames are captured and converted into character information by the Tesseract OCR engine. The input is the video frame, and the output is the extracted character data.
[0313] Step 5:
[0314] The server translates the transcribed audio data and extracted text data into the specified language using the Google Cloud Translation API. This involves applying a translation algorithm to the data and converting it to the specified target language. The input is text data, and the output is translated text data.
[0315] Step 6:
[0316] The server synthesizes the translated text data into speech using the Google Text-to-Speech API and generates an audio file. The input is the translated text data, and the output is the synthesized audio file.
[0317] Step 7:
[0318] The server integrates the translated audio and text into the original video data to generate a new video. Video editing software then edits the video to ensure that subtitles and audio are displayed correctly. The input consists of translated audio files and text data, while the output is the generated video data.
[0319] Step 8:
[0320] The server streams the generated video to the user's terminal in real time. The terminal synchronizes the received video data, making it viewable to the user without delay. The input is the generated video data, and the output is the video played on the terminal.
[0321] (Application Example 1)
[0322] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0323] In conventional systems, real-time multilingual translation of lectures and classes was difficult, creating a significant barrier to access for users in different language environments. Furthermore, the lack of visualization of audio and text information within videos limited the user's viewing experience. This hindered the equalization of education in a multicultural society.
[0324] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0325] In this invention, the server includes means for processing audio data in real time and translating it from a first language to a second language; means for extracting text information from video data and translating it from a first language to a second language; means for synchronizing the translation results with the video data and constructing generated video including the translated audio and text; communication means for delivering the generated video to a terminal in real time and enabling the user to specify a selective viewing language; and means for providing educational information in the user's native language in the selected language and displaying it via a visual device. This enables the user to overcome language barriers and view high-quality educational content in their native language in real time.
[0326] "Real-time" means that information processing is performed instantly on the spot.
[0327] "Audio data" refers to audio information represented in digital format.
[0328] The "primary language" is the language in which the original spoken or written text is expressed.
[0329] A "second language" is the language in which the translated information is intended to be expressed.
[0330] "Means of translation" refer to functions or processes for converting information from one language into another language.
[0331] "Video data" refers to digital data that includes visual information.
[0332] "Text information" refers to information expressed in the form of characters.
[0333] "Synchronizing" means making different pieces of information or actions coincide and match at the same time.
[0334] "Generated video" refers to video that has been newly created by processing the original video.
[0335] A "terminal" is an information processing device used by a user.
[0336] "To distribute" means to send information to a recipient.
[0337] "Communication methods" refer to the methods and technologies used to send and receive data.
[0338] "Auditory language" refers to language intended for users to understand through their sight and hearing.
[0339] "Educational information" refers to information that contains knowledge and content useful for learning.
[0340] A "visual device" is a hardware device that displays video information.
[0341] This invention describes a system for delivering educational content translated into multiple languages in real time using specific devices and software.
[0342] The server first converts the audio data into text using speech recognition technology. Cloud-based services such as Google Cloud Speech-to-Text or Amazon Transcribe can be used for speech recognition. This text is then translated into a second language by a translation engine. Multilingual translation services, such as the Google Translate API, are used in the translation process.
[0343] The server also uses optical character recognition (OCR) technology to extract text information from the video data. In this process, the text data on the whiteboard or blackboard is converted into a digital format and translated into the specified language. Using video processing software such as FFmpeg (described later), the translated text and audio are accurately overlaid onto the original video to construct the generated video.
[0344] The generated video is delivered to the user's device in real time. This device may be a smartphone, smart glasses, or a head-mounted display. The delivered video can be viewed in the user's chosen language, and playback is achieved with minimal latency.
[0345] For example, when a user watches a lecture given at a university in a different country, this system translates the lecture audio into the user's native language and also translates and displays the contents of the whiteboard, thereby providing equal educational opportunities that transcend language barriers.
[0346] An example of a prompt for a generative AI model might be: "Use a translation service to convert the lecture audio from English to Japanese in real time. Additionally, translate the text in the video into Japanese and display it in sync with the video."
[0347] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0348] Step 1:
[0349] The user uses their device to select the lecture or lesson content they wish to view and their preferred translation language, then submits a request. The user's selection information is used as input and sent to the server. The output consists of content identification information and the translation language setting.
[0350] Step 2:
[0351] The server retrieves the specified content via the network based on a request received from the user. The input is the user's request information, and the output is the audio and video data of the content. This process prepares the data to be viewed within the system.
[0352] Step 3:
[0353] The server uses speech recognition technology to convert the acquired audio data into text data. The input is audio data, and the output is text data converted from the audio. In this step, a service such as Google Cloud Speech-to-Text is used to convert the audio to text.
[0354] Step 4:
[0355] The server uses optical character recognition (OCR) technology to extract text information from video data. The input is video data, and the output is extracted text data. For example, OCR software such as Tesseract may be used.
[0356] Step 5:
[0357] The server translates text data into a specified second language. Input is text data obtained from audio and video, and output is translated text data. Real-time translation is performed using APIs such as Google Translate.
[0358] Step 6:
[0359] The server constructs a generated video by overlaying the translated audio and text data onto the original video data. The input is the translated audio and text along with the original video data, and the output is a viewable generated video. Video editing software such as FFmpeg is used to integrate the audio and text into the video.
[0360] Step 7:
[0361] The server streams the generated video to the user's terminal in real time. The input is the generated video data, and the output is the video played on the user's terminal. On the terminal side, the translated content is played back without delay.
[0362] Step 8:
[0363] Users can view lectures and lessons in their chosen language on their devices, enabling them to understand content in their native language. Users can also receive educational information in real time through streamed, generated video.
[0364] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0365] This invention provides a system that translates the audio and video content of lectures and classes into multiple languages in real time, and also has the function of recognizing the user's emotions. This system operates via a server, user terminals, an emotion engine, and a digital network connecting them.
[0366] The user's device is the one that sends viewing requests to the server. The user selects the lecture or lesson video they want to watch, specifies their desired translation language, and sends the request to the server.
[0367] The server receives a request from the user, retrieves the video stream over the network, and begins processing. Speech recognition technology is applied on the server, converting the video's audio track into text data. This text data is then translated into the specified language using generative AI.
[0368] For the visual elements of the video, the server uses optical character recognition technology to extract text information and translates it in the same way. The translation results are converted into speech in the specified language using speech synthesis technology and overlaid on the video.
[0369] Furthermore, the emotion engine analyzes the user's emotions in real time using data from the device's built-in camera or data provided by the user. This emotion data is sent to the server and influences the display of translated content. For example, if the emotion engine determines that the user is confused, the server can enlarge the explanatory text or slow down the playback of the audio commentary.
[0370] The user's device receives translated and emotion-adapted video streamed from the server and plays it back in real time. This allows users to view translated content in a way that is easier to understand and responds to their individual emotional state.
[0371] In this way, the present invention overcomes language barriers and improves the user experience, thereby providing a multicultural learning environment and realizing a system that enables users worldwide to access high-quality educational content.
[0372] The following describes the processing flow.
[0373] Step 1:
[0374] The user uses their device to select the lecture or lesson video they want to watch and sends a request to the server. This request includes the language in which they want the video translated.
[0375] Step 2:
[0376] The server receives a request from the user, connects to the video stream of the specified lecture or class, and prepares to retrieve the data.
[0377] Step 3:
[0378] The server acquires audio data from the video and uses speech recognition technology to convert this audio into text. The resulting text is then used in the translation process.
[0379] Step 4:
[0380] The server analyzes the video data and extracts visual text information from the video using optical character recognition technology. This text information is immediately subjected to a translation process.
[0381] Step 5:
[0382] The server translates audio text data and in-video text information into a second language specified by the user. This translation is optimized for real-time operation.
[0383] Step 6:
[0384] The server converts the translated text information into speech using speech synthesis technology and places it as subtitles on the video. This synthesized speech and text are synchronized with the original video.
[0385] Step 7:
[0386] The user's device captures their facial expressions with a camera, and an emotion engine analyzes the user's emotions in real time based on this data. The emotion engine then sends the analysis results to the server.
[0387] Step 8:
[0388] Based on data received from the emotion engine, the server adaptively adjusts the translated content according to the user's emotions. For example, if the user is confused, the server will deliver content that emphasizes explanatory parts or plays audio at a slower pace.
[0389] Step 9:
[0390] The server streams the adjusted and translated video to the user's device in real time.
[0391] Step 10:
[0392] The user's device receives translated video streamed from the server and plays it back while synchronizing the video and audio. This allows the user to instantly view content adjusted to their own language.
[0393] (Example 2)
[0394] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0395] It is necessary to remove language barriers in international lectures and classes and provide support to help participants gain a better understanding. Furthermore, there is a need to improve learning effectiveness by providing flexible content that responds to participants' emotional states.
[0396] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0397] In this invention, the server includes means for converting audio data into text data in real time using speech recognition technology and translating the text data from a first language to a second language; means for extracting character information from video data using optical character recognition technology and translating the extracted character information from a first language to a second language; means for converting the translation result into audio data using speech synthesis technology and constructing a generated video by overlaying the audio data and translated text onto video data; means for adjusting the display timing and method of the generated video based on emotion data obtained from a device that analyzes the user's emotions; and means for delivering the generated video to an information terminal in real time. As a result, the user can receive translated content adapted to their emotional state in real time.
[0398] "Speech recognition technology" is a technology that converts human speech into text data using a computer system.
[0399] "Optical character recognition technology" is a technology that extracts character information from image or video data.
[0400] "Translation means" refers to a process or device for converting text data into another language.
[0401] "Speech synthesis technology" is a technology that converts text data into speech data and outputs it as natural-sounding speech.
[0402] A "generative AI model" is an artificial intelligence system that performs natural language processing to generate and translate text data.
[0403] A "sentiment analysis device" is a device that analyzes a user's emotions in real time and acquires that data.
[0404] "Video data" refers to digital data that contains visual information.
[0405] An "information terminal" is an electronic device that transmits, receives, and processes data via a network.
[0406] This invention relates to a system that translates the audio and video content of lectures and classes into other languages in real time and also has the function of recognizing the user's emotions. This system consists of a server, information terminals, an emotion analysis device, and a network connecting them.
[0407] The user first uses an information terminal to select the content they want to view and specify the desired translation language. This request is sent to the server. The server uses speech recognition technology (e.g., commercial speech recognition software) to convert the acquired speech data into text data.
[0408] Next, a generative AI model (for example, a general-purpose generative AI engine) is used to translate the text data into the specified language. A prompt such as "Translate Japanese into English" is used.
[0409] For video data, the server extracts text information using optical character recognition technology (e.g., common OCR software) and translates this information. The translation result is converted into speech using speech synthesis technology (e.g., standard text-to-speech software) and incorporated into the video as an overlay.
[0410] For sentiment analysis, the camera and microphone built into the user's information terminal are used. The data acquired is analyzed in real time by a sentiment analysis device (e.g., a standard sentiment recognition API), and the sentiment information is sent to a server. Based on the sentiment information, the server can dynamically adjust how the generated video is displayed.
[0411] A concrete example is a user who wants to understand a Spanish physics lecture in English. The user selects the relevant Spanish lecture from their device and specifies English as the translation language. The audio and video are recognized and translated, and the audio and text are provided to the user in English in a synchronized format. Furthermore, the server can appropriately adjust the playback speed and level of detail according to the user's level of difficulty.
[0412] In this way, this system enables users to understand educational content in other languages in a way that is optimized for their own emotions.
[0413] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0414] Step 1:
[0415] The user selects the lecture or class content they wish to view using their information terminal and specifies their desired translation language. A request is generated containing the user's selected video data and desired language information. This request is sent to the server, where processing begins.
[0416] Step 2:
[0417] The server retrieves a specified video stream over the network. It uses video data based on user requests as input. The output is generated by converting audio data into text data using speech recognition technology. Specifically, the speech recognition engine analyzes the audio waveform and outputs it as text in the corresponding language.
[0418] Step 3:
[0419] The server uses a generative AI model to translate the text data obtained in step 2 into the specified language. The recognized original text data and a prompt message are used as input. A prompt message such as "Please translate Japanese into English" is used. The output is the translated text data.
[0420] Step 4:
[0421] The server extracts text information from video data using optical character recognition (OCR) technology. The input is video data, and the output is the extracted text information. The OCR engine analyzes the characters in the image and outputs them as corresponding text.
[0422] Step 5:
[0423] The character information extracted in Step 4 is similarly translated into the specified language using the generative AI model. The recognized character information and prompt text are used as input, and the output is the translated character information.
[0424] Step 6:
[0425] The server uses speech synthesis technology to convert translated text data into new speech data. The input is translated text data, and the output is speech data in the specified language. The speech synthesis engine analyzes the text and outputs it as synthesized speech.
[0426] Step 7:
[0427] The server overlays the generated audio data and translated text onto the video data to construct the generated video. The input consists of the generated audio data, translated text data, and original video data, and the output is the generated video in which they are integrated.
[0428] Step 8:
[0429] The emotion analysis device uses data obtained from the camera and microphone built into the device to detect the user's emotions and sends it to the server as temporary input data. This data indicates the emotional state, and a specific emotional pattern is obtained as output.
[0430] Step 9:
[0431] The server adjusts how the generated video is displayed based on the user's emotion data. The input is emotion data and generated video data, and the output is the generated video with the adjusted playback parameters.
[0432] Step 10:
[0433] The server delivers the adjusted, generated video to the information terminal in real time. Users view the content played on their terminal, receiving translations and emotionally-responsive learning support. The final output is the content played back in real time.
[0434] (Application Example 2)
[0435] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0436] In an increasingly globalized world, understanding educational and informational content provided in different languages remains challenging, transcending language barriers. Furthermore, the lack of content adjustments tailored to the audience's emotions and level of understanding leads to decreased learning efficiency. Therefore, a system is needed that provides real-time multilingual translation and customized information based on the audience's emotions.
[0437] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0438] In this invention, the server includes means for processing audio data in real time and translating it from a first language to a second language, means for extracting text information from video data and translating it from a first language to a second language, and means for acquiring the user's emotional state and adjusting the display method of the generated video based on the emotion. This enables the display of multilingual educational content in a customized manner according to the viewer's emotions, resulting in a more effective learning experience.
[0439] "Means for processing audio data in real time and translating from a first language to a second language" refers to a device or function that performs the process of instantly analyzing an audio signal, converting it into text, and then translating that text into a specified different language.
[0440] "Means for extracting text information from video data and translating it from a first language to a second language" refers to a device or function that identifies and acquires characters within a video and processes those characters to represent them in a different language.
[0441] "Means for synchronizing translation results with video data and constructing generated video including translated audio and text" refers to a device or function that performs the process of combining audio and text translated into different languages with the original video to create integrated video content.
[0442] "Means for delivering generated video to terminals in real time" refers to a network system and its control process that transfers the created translated video to the viewer's device without delay.
[0443] "Means for acquiring the user's emotional state and adjusting the display method of generated video based on that emotion" refers to a device or function that analyzes the user's facial expressions and behavioral data to recognize emotions and automatically changes the video playback speed, display content, etc., according to the analysis results.
[0444] The system for realizing this invention is composed of an integrated set of technologies, primarily consisting of servers, user terminals, and various software technologies.
[0445] First, the server receives a viewing request from the user. This request includes the content to be viewed and the desired translation language. The server acquires the audio data and converts it to text using speech recognition technology. This process primarily uses Python and speech recognition libraries.
[0446] Furthermore, the server uses optical character recognition (OCR) technology to extract text information from the video data. Simultaneously, a generative AI model is used to translate this text into the language specified by the user. At this stage, the OpenAI API is often used.
[0447] The server uses the user's device's camera to analyze facial expressions and behavior in real time in order to capture the user's emotional state. Machine learning libraries such as TensorFlow are used for emotion analysis.
[0448] The converted audio and text are synchronized with the video data as translation results. Speech synthesis technology is used here, and the resulting video is displayed in a way that adjusts based on the user's emotional state. For example, if the user is determined to be confused, the playback speed can be adjusted, or additional detailed explanations can be displayed. Finally, the adjusted video and audio are delivered to the device.
[0449] For example, if a Japanese user watching a technical lecture in English uses this system, and it determines that technical terms are difficult to understand, the system will display a detailed explanation of those terms in Japanese and play the audio commentary at a slower pace. The following is an example of a prompt to the generative AI model:
[0450] "Translate the following educational content from English to Japanese, emphasizing clarity for a high school student: English lesson content"
[0451] This makes it possible to provide a high-quality learning experience that transcends language and cultural barriers.
[0452] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0453] Step 1:
[0454] The server receives viewing requests from users. The input includes the ID of the content they want to view and the desired translation language. Based on this information, the server retrieves the corresponding video and audio data.
[0455] Step 2:
[0456] The server sends the acquired audio data to the speech recognition engine, which converts it into text data. The input is an audio signal, and the output is the recognized text. This process is performed by a Python speech recognition library.
[0457] Step 3:
[0458] The server extracts text information from video data using optical character recognition (OCR) technology. The input is video frames, and the output is the string data contained within the video. An OCR library is used in this step.
[0459] Step 4:
[0460] The server sends text data to a generation AI model for translation into the desired language. The input is text data in the original language, and the output is text data translated into the specified language. This translation process uses the OpenAI API.
[0461] Step 5:
[0462] The user's device captures facial expressions using its built-in camera and sends them to an emotion recognition engine. The input is facial image data, and the output is data indicating the user's emotional state. Libraries such as TensorFlow are used for emotion analysis.
[0463] Step 6:
[0464] The server synchronizes the translation results with the video data based on sentiment data and adjusts the display accordingly. The input consists of the translation results and sentiment data, and the output is a video with adjusted display settings. Specifically, this includes changes to font size and playback speed.
[0465] Step 7:
[0466] The server delivers generated, translated, and emotion-sensitive video to the terminal. The input is the adjusted video, and the output is this video displayed in real time on the user's terminal. This utilizes streaming technology over a network.
[0467] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0468] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0469] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0470] [Third Embodiment]
[0471] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0472] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0473] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0474] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0475] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0476] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0477] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0478] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0479] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0480] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0481] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0482] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0483] This invention provides a system that translates lectures and class videos into multiple languages in real time and delivers them to users. This system operates through a server, user terminals, and a digital network.
[0484] The user's device functions as a device that sends viewing requests to the server. The user selects the lecture or lesson video they want to watch and sends the request to the server. At this stage, they can specify their desired translation language.
[0485] The server receives requests from users and retrieves videos of selected lectures or classes via the network. Using speech recognition technology, the server analyzes the audio track of the video and converts it into text. This transcribed audio data is then translated in real time by the server.
[0486] Simultaneously, the server uses optical character recognition technology to extract text written on whiteboards or blackboards in the video. This visual text information is also translated into the selected language.
[0487] The translated audio and text are reconstructed on the server and overlaid onto the original video. For example, if a math lesson conducted in English is translated into Japanese, the English lecture audio is replaced with Japanese audio, and the mathematical formulas and explanations written on the whiteboard are translated into Japanese and displayed on the video. This process allows users to view learning content translated into their native language in real time.
[0488] The user's device functions as a device that receives translated video streamed from the server and plays it back in real time. On the device, the video and audio are synchronized and adjusted to flow without delay.
[0489] In this way, the present invention provides a system that transcends language barriers, offers a multicultural learning environment, and enables users worldwide to access high-quality educational content.
[0490] The following describes the processing flow.
[0491] Step 1:
[0492] Users send a request from their device to the server to view a video of a specific lecture or class. They also specify their preferred translation language at the same time.
[0493] Step 2:
[0494] The server receives a request from the user and connects to the specified video stream. It then prepares to retrieve the streamed video.
[0495] Step 3:
[0496] The server analyzes the audio track of the acquired video using speech recognition technology and converts the audio data into text. This transcribed audio data is then used in the translation process.
[0497] Step 4:
[0498] The server analyzes the video frames and uses optical character recognition technology to extract text information written on the whiteboard or blackboard. This information is also processed as text data.
[0499] Step 5:
[0500] The server translates both audio-text and video-text into a specified second language. This translation utilizes generative AI technology to achieve real-time processing.
[0501] Step 6:
[0502] The server uses speech synthesis technology to convert the translated audio text into speech in the specified language. The translated text is also synchronized with the original video, overlaid on top of it.
[0503] Step 7:
[0504] The server generates a video that integrates the translated audio and text, and streams it to the user's device in real time.
[0505] Step 8:
[0506] The user's device receives translated video transmitted from the server and plays it back in real time. Audio and video are synchronized, providing the user with a lag-free viewing experience.
[0507] (Example 1)
[0508] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0509] In modern society, multilingual translation is crucial for people who speak different languages to smoothly share and understand information. However, technologies for translating lectures and class videos in real time face many challenges, including accurate synchronization of audio and text, translation accuracy and speed, and smooth delivery to users. This invention aims to solve these problems and provide a system that promotes multicultural education and knowledge sharing by removing barriers between different languages.
[0510] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0511] In this invention, the server includes means for receiving requests to view lectures or classes from user terminals, means for acquiring video via the network based on the request and converting the audio track into text using speech recognition technology, and means for extracting text information from the video using optical character recognition technology. This enables the content of lectures and classes provided in different languages to be translated quickly and accurately into other languages, allowing users to view videos in their own language in real time.
[0512] A "user terminal" is a device used by a user to retrieve information from a server via the internet and to view video and audio.
[0513] A "server" is a computer system that receives requests from users on a network, processes the data, and provides the results to the user.
[0514] A "viewing request" is a request that a user sends to a server to view a video of a specific lecture or class.
[0515] A "network" is a system that connects computers and devices to each other, enabling the transmission and reception of data.
[0516] "Speech recognition technology" is a general term for the processes and technologies used to convert speech data into text data.
[0517] Optical Character Recognition (OCR) is a technology that extracts character information from image data, and is often abbreviated as OCR.
[0518] "Multilingual translation technology" is a system of technologies for converting text or audio in one language into another language.
[0519] "Speech synthesis technology" is a technology that converts text into audio data and reproduces it audibly.
[0520] "Means of generating video" refers to the process of integrating translated audio and text to create new video data.
[0521] "Means of real-time delivery" refers to technologies and methods that enable the immediate transmission of generated video to the user's device, allowing for viewing without delay.
[0522] This invention is a system that translates lectures and class videos into multiple languages in real time and provides them to users. The system consists of a server, user terminals, and a digital network connecting them.
[0523] The server is a computer system that receives viewing requests from users, which include a video ID and the desired translation language. Upon recognizing the request, the server retrieves the relevant video metadata from the database and pulls the video file from the appropriate storage system. For speech recognition technology, the Google Speech-to-Text API can be used to convert the audio track into text data.
[0524] Furthermore, Tesseract OCR is applied to optical character recognition technology to extract text information from the video. This text is then translated in real time into the specified language using multilingual translation technologies such as the Google Cloud Translation API. The translated text is then generated as an audio file using speech synthesis technology (e.g., Google Text-to-Speech API).
[0525] The user terminal receives translated video streamed from the server and provides the video to the user through sight and sound. The terminal functions as a media player, decoding and buffering the received data and adjusting it so that the video and audio are played in sync in real time.
[0526] For example, if a history lesson conducted in English is translated into French, and the user requests French on their device, the server can retrieve the English lecture video, translate the audio and text data into French in real time, and provide it to the user as a streaming video. An example of a prompt related to this process would be, "Please provide a real-time translation of an English history lecture into French, including both the spoken lecture and any written content on blackboards, making it viewable in a video format."
[0527] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0528] Step 1:
[0529] The user uses their device to request to view a video of a lecture or class. Specifically, they enter the video title and desired translation language into the user interface and press the "Submit Request" button. At this time, the video ID and translation language data are sent to the server as input. The output is the request data received by the server.
[0530] Step 2:
[0531] The server parses the request received from the user and retrieves metadata from the database based on the video ID. This data includes links to the necessary video files based on the request. The input is the request data, and the output is the retrieved video metadata.
[0532] Step 3:
[0533] The server downloads the relevant video file from the storage system. This video file contains an audio track, which is converted to text by a speech recognition system. Specifically, the Google Speech-to-Text API receives the audio data and generates the corresponding text data. The input is the video file, and the output is the transcribed audio data.
[0534] Step 4:
[0535] The server extracts text information from video frames using optical character recognition (OCR) technology. Video frames are captured and converted into text information by the Tesseract OCR engine. The input is the video frame, and the output is the extracted text data.
[0536] Step 5:
[0537] The server translates the transcribed audio data and extracted text data into the specified language using the Google Cloud Translation API. This involves applying a translation algorithm to the data and converting it to the specified target language. The input is text data, and the output is translated text data.
[0538] Step 6:
[0539] The server synthesizes the translated text data into speech using the Google Text-to-Speech API and generates an audio file. The input is the translated text data, and the output is the synthesized audio file.
[0540] Step 7:
[0541] The server integrates the translated audio and text into the original video data to generate a new video. Video editing software then edits the video to ensure that subtitles and audio are displayed correctly. The input consists of translated audio files and text data, while the output is the generated video data.
[0542] Step 8:
[0543] The server streams the generated video to the user's terminal in real time. The terminal synchronizes the received video data, making it viewable to the user without delay. The input is the generated video data, and the output is the video played on the terminal.
[0544] (Application Example 1)
[0545] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0546] In conventional systems, real-time multilingual translation of lectures and classes was difficult, creating a significant barrier to access for users in different language environments. Furthermore, the lack of visualization of audio and text information within videos limited the user's viewing experience. This hindered the equalization of education in a multicultural society.
[0547] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0548] In this invention, the server includes means for processing audio data in real time and translating it from a first language to a second language; means for extracting text information from video data and translating it from a first language to a second language; means for synchronizing the translation results with the video data and constructing generated video including the translated audio and text; communication means for delivering the generated video to a terminal in real time and enabling the user to specify a selective viewing language; and means for providing educational information in the user's native language in the selected language and displaying it via a visual device. This enables the user to overcome language barriers and view high-quality educational content in their native language in real time.
[0549] "Real-time" means that information processing is performed instantly on the spot.
[0550] "Audio data" refers to audio information represented in digital format.
[0551] The "primary language" is the language in which the original spoken or written text is expressed.
[0552] A "second language" is the language in which the translated information is intended to be expressed.
[0553] "Means of translation" refer to functions or processes for converting information from one language into another language.
[0554] "Video data" refers to digital data that includes visual information.
[0555] "Text information" refers to information expressed in the form of characters.
[0556] "Synchronizing" means making different pieces of information or actions coincide and match at the same time.
[0557] "Generated video" refers to video that has been newly created by processing the original video.
[0558] A "terminal" is an information processing device used by a user.
[0559] "To distribute" means to send information to a recipient.
[0560] "Communication methods" refer to the methods and technologies used to send and receive data.
[0561] "Auditory language" refers to language intended for users to understand through their sight and hearing.
[0562] "Educational information" refers to information that contains knowledge and content useful for learning.
[0563] A "visual device" is a hardware device that displays video information.
[0564] This invention describes a system for delivering educational content translated into multiple languages in real time using specific devices and software.
[0565] The server first converts the audio data into text using speech recognition technology. Cloud-based services such as Google Cloud Speech-to-Text or Amazon Transcribe can be used for speech recognition. This text is then translated into a second language by a translation engine. Multilingual translation services, such as the Google Translate API, are used in the translation process.
[0566] The server also uses optical character recognition (OCR) technology to extract text information from the video data. In this process, the text data on the whiteboard or blackboard is converted into a digital format and translated into the specified language. Using video processing software such as FFmpeg (described later), the translated text and audio are accurately overlaid onto the original video to construct the generated video.
[0567] The generated video is delivered to the user's device in real time. This device may be a smartphone, smart glasses, or a head-mounted display. The delivered video can be viewed in the user's chosen language, and playback is achieved with minimal latency.
[0568] For example, when a user watches a lecture given at a university in a different country, this system translates the lecture audio into the user's native language and also translates and displays the contents of the whiteboard, thereby providing equal educational opportunities that transcend language barriers.
[0569] An example of a prompt for a generative AI model might be: "Use a translation service to convert the lecture audio from English to Japanese in real time. Additionally, translate the text in the video into Japanese and display it in sync with the video."
[0570] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0571] Step 1:
[0572] The user uses their device to select the lecture or lesson content they wish to view and their preferred translation language, then submits a request. The user's selection information is used as input and sent to the server. The output consists of content identification information and the translation language setting.
[0573] Step 2:
[0574] The server retrieves the specified content via the network based on a request received from the user. The input is the user's request information, and the output is the audio and video data of the content. This process prepares the data to be viewed within the system.
[0575] Step 3:
[0576] The server uses speech recognition technology to convert the acquired audio data into text data. The input is audio data, and the output is text data converted from the audio. In this step, a service such as Google Cloud Speech-to-Text is used to convert the audio to text.
[0577] Step 4:
[0578] The server uses optical character recognition (OCR) technology to extract text information from video data. The input is video data, and the output is extracted text data. For example, OCR software such as Tesseract may be used.
[0579] Step 5:
[0580] The server translates text data into a specified second language. Input is text data obtained from audio and video, and output is translated text data. Real-time translation is performed using APIs such as Google Translate.
[0581] Step 6:
[0582] The server constructs a generated video by overlaying the translated audio and text data onto the original video data. The input is the translated audio and text along with the original video data, and the output is a viewable generated video. Video editing software such as FFmpeg is used to integrate the audio and text into the video.
[0583] Step 7:
[0584] The server streams the generated video to the user's terminal in real time. The input is the generated video data, and the output is the video played on the user's terminal. On the terminal side, the translated content is played back without delay.
[0585] Step 8:
[0586] Users can view lectures and lessons in their chosen language on their devices, enabling them to understand content in their native language. Users can also receive educational information in real time through streamed, generated video.
[0587] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0588] This invention provides a system that translates the audio and video content of lectures and classes into multiple languages in real time, and also has the function of recognizing the user's emotions. This system operates via a server, user terminals, an emotion engine, and a digital network connecting them.
[0589] The user's device is the one that sends viewing requests to the server. The user selects the lecture or lesson video they want to watch, specifies their desired translation language, and sends the request to the server.
[0590] The server receives a request from the user, retrieves the video stream over the network, and begins processing. Speech recognition technology is applied on the server, converting the video's audio track into text data. This text data is then translated into the specified language using generative AI.
[0591] For the visual elements of the video, the server uses optical character recognition technology to extract text information and translates it in the same way. The translation results are converted into speech in the specified language using speech synthesis technology and overlaid on the video.
[0592] Furthermore, the emotion engine analyzes the user's emotions in real time using data from the device's built-in camera or data provided by the user. This emotion data is sent to the server and influences the display of translated content. For example, if the emotion engine determines that the user is confused, the server can enlarge the explanatory text or slow down the playback of the audio commentary.
[0593] The user's device receives translated and emotion-adapted video streamed from the server and plays it back in real time. This allows users to view translated content in a way that is easier to understand and responds to their individual emotional state.
[0594] In this way, the present invention overcomes language barriers and improves the user experience, thereby providing a multicultural learning environment and realizing a system that enables users worldwide to access high-quality educational content.
[0595] The following describes the processing flow.
[0596] Step 1:
[0597] The user uses their device to select the lecture or lesson video they want to watch and sends a request to the server. This request includes the language in which they want the video translated.
[0598] Step 2:
[0599] The server receives a request from the user, connects to the video stream of the specified lecture or class, and prepares to retrieve the data.
[0600] Step 3:
[0601] The server acquires audio data from the video and uses speech recognition technology to convert this audio into text. The resulting text is then used in the translation process.
[0602] Step 4:
[0603] The server analyzes the video data and extracts visual text information from the video using optical character recognition technology. This text information is immediately subjected to a translation process.
[0604] Step 5:
[0605] The server translates audio text data and in-video text information into a second language specified by the user. This translation is optimized for real-time operation.
[0606] Step 6:
[0607] The server converts the translated text information into speech using speech synthesis technology and places it as subtitles on the video. This synthesized speech and text are synchronized with the original video.
[0608] Step 7:
[0609] The user's device captures their facial expressions with a camera, and an emotion engine analyzes the user's emotions in real time based on this data. The emotion engine then sends the analysis results to the server.
[0610] Step 8:
[0611] Based on data received from the emotion engine, the server adaptively adjusts the translated content according to the user's emotions. For example, if the user is confused, the server will deliver content that emphasizes explanatory parts or plays audio at a slower pace.
[0612] Step 9:
[0613] The server streams the adjusted and translated video to the user's device in real time.
[0614] Step 10:
[0615] The user's device receives translated video streamed from the server and plays it back while synchronizing the video and audio. This allows the user to instantly view content adjusted to their own language.
[0616] (Example 2)
[0617] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0618] It is necessary to remove language barriers in international lectures and classes and provide support to help participants gain a better understanding. Furthermore, there is a need to improve learning effectiveness by providing flexible content that responds to participants' emotional states.
[0619] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0620] In this invention, the server includes means for converting audio data into text data in real time using speech recognition technology and translating the text data from a first language to a second language; means for extracting character information from video data using optical character recognition technology and translating the extracted character information from a first language to a second language; means for converting the translation result into audio data using speech synthesis technology and constructing a generated video by overlaying the audio data and translated text onto video data; means for adjusting the display timing and method of the generated video based on emotion data obtained from a device that analyzes the user's emotions; and means for delivering the generated video to an information terminal in real time. As a result, the user can receive translated content adapted to their emotional state in real time.
[0621] "Speech recognition technology" is a technology that converts human speech into text data using a computer system.
[0622] "Optical character recognition technology" is a technology that extracts character information from image or video data.
[0623] "Translation means" refers to a process or device for converting text data into another language.
[0624] "Speech synthesis technology" is a technology that converts text data into speech data and outputs it as natural-sounding speech.
[0625] A "generative AI model" is an artificial intelligence system that performs natural language processing to generate and translate text data.
[0626] A "sentiment analysis device" is a device that analyzes a user's emotions in real time and acquires that data.
[0627] "Video data" refers to digital data that contains visual information.
[0628] An "information terminal" is an electronic device that transmits, receives, and processes data via a network.
[0629] This invention relates to a system that translates the audio and video content of lectures and classes into other languages in real time and also has the function of recognizing the user's emotions. This system consists of a server, information terminals, an emotion analysis device, and a network connecting them.
[0630] The user first uses an information terminal to select the content they want to view and specify the desired translation language. This request is sent to the server. The server uses speech recognition technology (e.g., commercial speech recognition software) to convert the acquired speech data into text data.
[0631] Next, a generative AI model (for example, a general-purpose generative AI engine) is used to translate the text data into the specified language. A prompt such as "Translate Japanese into English" is used.
[0632] For video data, the server extracts text information using optical character recognition technology (e.g., common OCR software) and translates this information. The translation result is converted into speech using speech synthesis technology (e.g., standard text-to-speech software) and incorporated into the video as an overlay.
[0633] For sentiment analysis, the camera and microphone built into the user's information terminal are used. The data acquired is analyzed in real time by a sentiment analysis device (e.g., a standard sentiment recognition API), and the sentiment information is sent to a server. Based on the sentiment information, the server can dynamically adjust how the generated video is displayed.
[0634] A concrete example is a user who wants to understand a Spanish physics lecture in English. The user selects the relevant Spanish lecture from their device and specifies English as the translation language. The audio and video are recognized and translated, and the audio and text are provided to the user in English in a synchronized format. Furthermore, the server can appropriately adjust the playback speed and level of detail according to the user's level of difficulty.
[0635] In this way, this system enables users to understand educational content in other languages in a way that is optimized for their own emotions.
[0636] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0637] Step 1:
[0638] The user selects the lecture or class content they wish to view using their information terminal and specifies their desired translation language. A request is generated containing the user's selected video data and desired language information. This request is sent to the server, where processing begins.
[0639] Step 2:
[0640] The server retrieves a specified video stream over the network. It uses video data based on user requests as input. The output is generated by converting audio data into text data using speech recognition technology. Specifically, the speech recognition engine analyzes the audio waveform and outputs it as text in the corresponding language.
[0641] Step 3:
[0642] The server uses a generative AI model to translate the text data obtained in step 2 into the specified language. The recognized original text data and a prompt message are used as input. A prompt message such as "Please translate Japanese into English" is used. The output is the translated text data.
[0643] Step 4:
[0644] The server extracts text information from video data using optical character recognition (OCR) technology. The input is video data, and the output is the extracted text information. The OCR engine analyzes the characters in the image and outputs them as corresponding text.
[0645] Step 5:
[0646] The character information extracted in Step 4 is similarly translated into the specified language using the generative AI model. The recognized character information and prompt text are used as input, and the output is the translated character information.
[0647] Step 6:
[0648] The server uses speech synthesis technology to convert translated text data into new speech data. The input is translated text data, and the output is speech data in the specified language. The speech synthesis engine analyzes the text and outputs it as synthesized speech.
[0649] Step 7:
[0650] The server overlays the generated audio data and translated text onto the video data to construct the generated video. The input consists of the generated audio data, translated text data, and original video data, and the output is the generated video in which they are integrated.
[0651] Step 8:
[0652] The emotion analysis device uses data obtained from the camera and microphone built into the device to detect the user's emotions and sends it to the server as temporary input data. This data indicates the emotional state, and a specific emotional pattern is obtained as output.
[0653] Step 9:
[0654] The server adjusts how the generated video is displayed based on the user's emotion data. The input is emotion data and generated video data, and the output is the generated video with the adjusted playback parameters.
[0655] Step 10:
[0656] The server delivers the adjusted, generated video to the information terminal in real time. Users view the content played on their terminal, receiving translations and emotionally-responsive learning support. The final output is the content played back in real time.
[0657] (Application Example 2)
[0658] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0659] In an increasingly globalized world, understanding educational and informational content provided in different languages remains challenging, transcending language barriers. Furthermore, the lack of content adjustments tailored to the audience's emotions and level of understanding leads to decreased learning efficiency. Therefore, a system is needed that provides real-time multilingual translation and customized information based on the audience's emotions.
[0660] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0661] In this invention, the server includes means for processing audio data in real time and translating it from a first language to a second language, means for extracting text information from video data and translating it from a first language to a second language, and means for acquiring the user's emotional state and adjusting the display method of the generated video based on the emotion. This enables the display of multilingual educational content in a customized manner according to the viewer's emotions, resulting in a more effective learning experience.
[0662] "Means for processing audio data in real time and translating from a first language to a second language" refers to a device or function that performs the process of instantly analyzing an audio signal, converting it into text, and then translating that text into a specified different language.
[0663] "Means for extracting text information from video data and translating it from a first language to a second language" refers to a device or function that identifies and acquires characters within a video and processes those characters to represent them in a different language.
[0664] "Means for synchronizing translation results with video data and constructing generated video including translated audio and text" refers to a device or function that performs the process of combining audio and text translated into different languages with the original video to create integrated video content.
[0665] "Means for delivering generated video to terminals in real time" refers to a network system and its control process that transfers the created translated video to the viewer's device without delay.
[0666] "Means for acquiring the user's emotional state and adjusting the display method of generated video based on that emotion" refers to a device or function that analyzes the user's facial expressions and behavioral data to recognize emotions and automatically changes the video playback speed, display content, etc., according to the analysis results.
[0667] The system for realizing this invention is composed of an integrated set of technologies, primarily consisting of servers, user terminals, and various software technologies.
[0668] First, the server receives a viewing request from the user. This request includes the content to be viewed and the desired translation language. The server acquires the audio data and converts it to text using speech recognition technology. This process primarily uses Python and speech recognition libraries.
[0669] Furthermore, the server uses optical character recognition (OCR) technology to extract text information from the video data. Simultaneously, a generative AI model is used to translate this text into the language specified by the user. At this stage, the OpenAI API is often used.
[0670] The server uses the user's device's camera to analyze facial expressions and behavior in real time in order to capture the user's emotional state. Machine learning libraries such as TensorFlow are used for emotion analysis.
[0671] The converted audio and text are synchronized with the video data as translation results. Speech synthesis technology is used here, and the resulting video is displayed in a way that adjusts based on the user's emotional state. For example, if the user is determined to be confused, the playback speed can be adjusted, or additional detailed explanations can be displayed. Finally, the adjusted video and audio are delivered to the device.
[0672] For example, if a Japanese user watching a technical lecture in English uses this system, and it determines that technical terms are difficult to understand, the system will display a detailed explanation of those terms in Japanese and play the audio commentary at a slower pace. The following is an example of a prompt to the generative AI model:
[0673] "Translate the following educational content from English to Japanese, emphasizing clarity for a high school student: English lesson content"
[0674] This makes it possible to provide a high-quality learning experience that transcends language and cultural barriers.
[0675] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0676] Step 1:
[0677] The server receives viewing requests from users. The input includes the ID of the content they want to view and the desired translation language. Based on this information, the server retrieves the corresponding video and audio data.
[0678] Step 2:
[0679] The server sends the acquired audio data to the speech recognition engine, which converts it into text data. The input is an audio signal, and the output is the recognized text. This process is performed by a Python speech recognition library.
[0680] Step 3:
[0681] The server extracts text information from video data using optical character recognition (OCR) technology. The input is video frames, and the output is the string data contained within the video. An OCR library is used in this step.
[0682] Step 4:
[0683] The server sends text data to a generation AI model for translation into the desired language. The input is text data in the original language, and the output is text data translated into the specified language. This translation process uses the OpenAI API.
[0684] Step 5:
[0685] The user's device captures facial expressions using its built-in camera and sends them to an emotion recognition engine. The input is facial image data, and the output is data indicating the user's emotional state. Libraries such as TensorFlow are used for emotion analysis.
[0686] Step 6:
[0687] The server synchronizes the translation results with the video data based on sentiment data and adjusts the display accordingly. The input consists of the translation results and sentiment data, and the output is a video with adjusted display settings. Specifically, this includes changes to font size and playback speed.
[0688] Step 7:
[0689] The server delivers generated, translated, and emotion-sensitive video to the terminal. The input is the adjusted video, and the output is this video displayed in real time on the user's terminal. This utilizes streaming technology over a network.
[0690] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0691] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0692] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0693] [Fourth Embodiment]
[0694] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0695] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0696] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0697] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0698] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0699] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0700] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0701] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0702] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0703] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0704] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0705] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0706] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0707] This invention provides a system that translates lectures and class videos into multiple languages in real time and delivers them to users. This system operates through a server, user terminals, and a digital network.
[0708] The user's device functions as a device that sends viewing requests to the server. The user selects the lecture or lesson video they want to watch and sends the request to the server. At this stage, they can specify their desired translation language.
[0709] The server receives requests from users and retrieves videos of selected lectures or classes via the network. Using speech recognition technology, the server analyzes the audio track of the video and converts it into text. This transcribed audio data is then translated in real time by the server.
[0710] Simultaneously, the server uses optical character recognition technology to extract text written on whiteboards or blackboards in the video. This visual text information is also translated into the selected language.
[0711] The translated audio and text are reconstructed on the server and overlaid onto the original video. For example, if a math lesson conducted in English is translated into Japanese, the English lecture audio is replaced with Japanese audio, and the mathematical formulas and explanations written on the whiteboard are translated into Japanese and displayed on the video. This process allows users to view learning content translated into their native language in real time.
[0712] The user's device functions as a device that receives translated video streamed from the server and plays it back in real time. On the device, the video and audio are synchronized and adjusted to flow without delay.
[0713] In this way, the present invention provides a system that transcends language barriers, offers a multicultural learning environment, and enables users worldwide to access high-quality educational content.
[0714] The following describes the processing flow.
[0715] Step 1:
[0716] Users send a request from their device to the server to view a video of a specific lecture or class. They also specify their preferred translation language at the same time.
[0717] Step 2:
[0718] The server receives a request from the user and connects to the specified video stream. It then prepares to retrieve the streamed video.
[0719] Step 3:
[0720] The server analyzes the audio track of the acquired video using speech recognition technology and converts the audio data into text. This transcribed audio data is then used in the translation process.
[0721] Step 4:
[0722] The server analyzes the video frames and uses optical character recognition technology to extract text information written on the whiteboard or blackboard. This information is also processed as text data.
[0723] Step 5:
[0724] The server translates both audio-text and video-text into a specified second language. This translation utilizes generative AI technology to achieve real-time processing.
[0725] Step 6:
[0726] The server uses speech synthesis technology to convert the translated audio text into speech in the specified language. The translated text is also synchronized with the original video, overlaid on top of it.
[0727] Step 7:
[0728] The server generates a video that integrates the translated audio and text, and streams it to the user's device in real time.
[0729] Step 8:
[0730] The user's device receives translated video transmitted from the server and plays it back in real time. Audio and video are synchronized, providing the user with a lag-free viewing experience.
[0731] (Example 1)
[0732] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0733] In modern society, multilingual translation is crucial for people who speak different languages to smoothly share and understand information. However, technologies for translating lectures and class videos in real time face many challenges, including accurate synchronization of audio and text, translation accuracy and speed, and smooth delivery to users. This invention aims to solve these problems and provide a system that promotes multicultural education and knowledge sharing by removing barriers between different languages.
[0734] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0735] In this invention, the server includes means for receiving requests to view lectures or classes from user terminals, means for acquiring video via the network based on the request and converting the audio track into text using speech recognition technology, and means for extracting text information from the video using optical character recognition technology. This enables the content of lectures and classes provided in different languages to be translated quickly and accurately into other languages, allowing users to view videos in their own language in real time.
[0736] A "user terminal" is a device used by a user to retrieve information from a server via the internet and to view video and audio.
[0737] A "server" is a computer system that receives requests from users on a network, processes the data, and provides the results to the user.
[0738] A "viewing request" is a request that a user sends to a server to view a video of a specific lecture or class.
[0739] A "network" is a system that connects computers and devices to each other, enabling the transmission and reception of data.
[0740] "Speech recognition technology" is a general term for the processes and technologies used to convert speech data into text data.
[0741] Optical Character Recognition (OCR) is a technology that extracts character information from image data, and is often abbreviated as OCR.
[0742] "Multilingual translation technology" is a system of technologies for converting text or audio in one language into another language.
[0743] "Speech synthesis technology" is a technology that converts text into audio data and reproduces it audibly.
[0744] "Means of generating video" refers to the process of integrating translated audio and text to create new video data.
[0745] "Means of real-time delivery" refers to technologies and methods that enable the immediate transmission of generated video to the user's device, allowing for viewing without delay.
[0746] This invention is a system that translates lectures and class videos into multiple languages in real time and provides them to users. The system consists of a server, user terminals, and a digital network connecting them.
[0747] The server is a computer system that receives viewing requests from users, which include a video ID and the desired translation language. Upon recognizing the request, the server retrieves the relevant video metadata from the database and pulls the video file from the appropriate storage system. For speech recognition technology, the Google Speech-to-Text API can be used to convert the audio track into text data.
[0748] Furthermore, Tesseract OCR is applied to optical character recognition technology to extract text information from the video. This text is then translated in real time into the specified language using multilingual translation technologies such as the Google Cloud Translation API. The translated text is then generated as an audio file using speech synthesis technology (e.g., Google Text-to-Speech API).
[0749] The user terminal receives translated video streamed from the server and provides the video to the user through sight and sound. The terminal functions as a media player, decoding and buffering the received data and adjusting it so that the video and audio are played in sync in real time.
[0750] For example, if a history lesson conducted in English is translated into French, and the user requests French on their device, the server can retrieve the English lecture video, translate the audio and text data into French in real time, and provide it to the user as a streaming video. An example of a prompt related to this process would be, "Please provide a real-time translation of an English history lecture into French, including both the spoken lecture and any written content on blackboards, making it viewable in a video format."
[0751] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0752] Step 1:
[0753] The user uses their device to request to view a video of a lecture or class. Specifically, they enter the video title and desired translation language into the user interface and press the "Submit Request" button. At this time, the video ID and translation language data are sent to the server as input. The output is the request data received by the server.
[0754] Step 2:
[0755] The server parses the request received from the user and retrieves metadata from the database based on the video ID. This data includes links to the necessary video files based on the request. The input is the request data, and the output is the retrieved video metadata.
[0756] Step 3:
[0757] The server downloads the relevant video file from the storage system. This video file contains an audio track, which is converted to text by a speech recognition system. Specifically, the Google Speech-to-Text API receives the audio data and generates the corresponding text data. The input is the video file, and the output is the transcribed audio data.
[0758] Step 4:
[0759] The server extracts text information from video frames using optical character recognition (OCR) technology. Video frames are captured and converted into text information by the Tesseract OCR engine. The input is the video frame, and the output is the extracted text data.
[0760] Step 5:
[0761] The server translates the transcribed audio data and extracted text data into the specified language using the Google Cloud Translation API. This involves applying a translation algorithm to the data and converting it to the specified target language. The input is text data, and the output is translated text data.
[0762] Step 6:
[0763] The server synthesizes the translated text data into speech using the Google Text-to-Speech API and generates an audio file. The input is the translated text data, and the output is the synthesized audio file.
[0764] Step 7:
[0765] The server integrates the translated audio and text into the original video data to generate a new video. Video editing software then edits the video to ensure that subtitles and audio are displayed correctly. The input consists of translated audio files and text data, while the output is the generated video data.
[0766] Step 8:
[0767] The server streams the generated video to the user's terminal in real time. The terminal synchronizes the received video data, making it viewable to the user without delay. The input is the generated video data, and the output is the video played on the terminal.
[0768] (Application Example 1)
[0769] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0770] In conventional systems, real-time multilingual translation of lectures and classes was difficult, creating a significant barrier to access for users in different language environments. Furthermore, the lack of visualization of audio and text information within videos limited the user's viewing experience. This hindered the equalization of education in a multicultural society.
[0771] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0772] In this invention, the server includes means for processing audio data in real time and translating it from a first language to a second language; means for extracting text information from video data and translating it from a first language to a second language; means for synchronizing the translation results with the video data and constructing generated video including the translated audio and text; communication means for delivering the generated video to a terminal in real time and enabling the user to specify a selective viewing language; and means for providing educational information in the user's native language in the selected language and displaying it via a visual device. This enables the user to overcome language barriers and view high-quality educational content in their native language in real time.
[0773] "Real-time" means that information processing is performed instantly on the spot.
[0774] "Audio data" refers to audio information represented in digital format.
[0775] The "primary language" is the language in which the original spoken or written text is expressed.
[0776] A "second language" is the language in which the translated information is intended to be expressed.
[0777] "Means of translation" refer to functions or processes for converting information from one language into another language.
[0778] "Video data" refers to digital data that includes visual information.
[0779] "Text information" refers to information expressed in the form of characters.
[0780] "Synchronizing" means making different pieces of information or actions coincide and match at the same time.
[0781] "Generated video" refers to video that has been newly created by processing the original video.
[0782] A "terminal" is an information processing device used by a user.
[0783] "To distribute" means to send information to a recipient.
[0784] "Communication methods" refer to the methods and technologies used to send and receive data.
[0785] "Auditory language" refers to language intended for users to understand through their sight and hearing.
[0786] "Educational information" refers to information that contains knowledge and content useful for learning.
[0787] A "visual device" is a hardware device that displays video information.
[0788] This invention describes a system for delivering educational content translated into multiple languages in real time using specific devices and software.
[0789] The server first converts the audio data into text using speech recognition technology. Cloud-based services such as Google Cloud Speech-to-Text or Amazon Transcribe can be used for speech recognition. This text is then translated into a second language by a translation engine. Multilingual translation services, such as the Google Translate API, are used in the translation process.
[0790] The server also uses optical character recognition (OCR) technology to extract text information from the video data. In this process, the text data on the whiteboard or blackboard is converted into a digital format and translated into the specified language. Using video processing software such as FFmpeg (described later), the translated text and audio are accurately overlaid onto the original video to construct the generated video.
[0791] The generated video is delivered to the user's device in real time. This device may be a smartphone, smart glasses, or a head-mounted display. The delivered video can be viewed in the user's chosen language, and playback is achieved with minimal latency.
[0792] For example, when a user watches a lecture given at a university in a different country, this system translates the lecture audio into the user's native language and also translates and displays the contents of the whiteboard, thereby providing equal educational opportunities that transcend language barriers.
[0793] An example of a prompt for a generative AI model might be: "Use a translation service to convert the lecture audio from English to Japanese in real time. Additionally, translate the text in the video into Japanese and display it in sync with the video."
[0794] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0795] Step 1:
[0796] The user uses their device to select the lecture or lesson content they wish to view and their preferred translation language, then submits a request. The user's selection information is used as input and sent to the server. The output consists of content identification information and the translation language setting.
[0797] Step 2:
[0798] The server retrieves the specified content via the network based on a request received from the user. The input is the user's request information, and the output is the audio and video data of the content. This process prepares the data to be viewed within the system.
[0799] Step 3:
[0800] The server uses speech recognition technology to convert the acquired audio data into text data. The input is audio data, and the output is text data converted from the audio. In this step, a service such as Google Cloud Speech-to-Text is used to convert the audio to text.
[0801] Step 4:
[0802] The server uses optical character recognition (OCR) technology to extract text information from video data. The input is video data, and the output is extracted text data. For example, OCR software such as Tesseract may be used.
[0803] Step 5:
[0804] The server translates text data into a specified second language. Input is text data obtained from audio and video, and output is translated text data. Real-time translation is performed using APIs such as Google Translate.
[0805] Step 6:
[0806] The server constructs a generated video by overlaying the translated audio and text data onto the original video data. The input is the translated audio and text along with the original video data, and the output is a viewable generated video. Video editing software such as FFmpeg is used to integrate the audio and text into the video.
[0807] Step 7:
[0808] The server streams the generated video to the user's terminal in real time. The input is the generated video data, and the output is the video played on the user's terminal. On the terminal side, the translated content is played back without delay.
[0809] Step 8:
[0810] Users can view lectures and lessons in their chosen language on their devices, enabling them to understand content in their native language. Users can also receive educational information in real time through streamed, generated video.
[0811] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0812] This invention provides a system that translates the audio and video content of lectures and classes into multiple languages in real time, and also has the function of recognizing the user's emotions. This system operates via a server, user terminals, an emotion engine, and a digital network connecting them.
[0813] The user's device is the one that sends viewing requests to the server. The user selects the lecture or lesson video they want to watch, specifies their desired translation language, and sends the request to the server.
[0814] The server receives a request from the user, retrieves the video stream over the network, and begins processing. Speech recognition technology is applied on the server, converting the video's audio track into text data. This text data is then translated into the specified language using generative AI.
[0815] For the visual elements of the video, the server uses optical character recognition technology to extract text information and translates it in the same way. The translation results are converted into speech in the specified language using speech synthesis technology and overlaid on the video.
[0816] Furthermore, the emotion engine analyzes the user's emotions in real time using data from the device's built-in camera or data provided by the user. This emotion data is sent to the server and influences the display of translated content. For example, if the emotion engine determines that the user is confused, the server can enlarge the explanatory text or slow down the playback of the audio commentary.
[0817] The user's device receives translated and emotion-adapted video streamed from the server and plays it back in real time. This allows users to view translated content in a way that is easier to understand and responds to their individual emotional state.
[0818] In this way, the present invention overcomes language barriers and improves the user experience, thereby providing a multicultural learning environment and realizing a system that enables users worldwide to access high-quality educational content.
[0819] The following describes the processing flow.
[0820] Step 1:
[0821] The user uses their device to select the lecture or lesson video they want to watch and sends a request to the server. This request includes the language in which they want the video translated.
[0822] Step 2:
[0823] The server receives a request from the user, connects to the video stream of the specified lecture or class, and prepares to retrieve the data.
[0824] Step 3:
[0825] The server acquires audio data from the video and uses speech recognition technology to convert this audio into text. The resulting text is then used in the translation process.
[0826] Step 4:
[0827] The server analyzes the video data and extracts visual text information from the video using optical character recognition technology. This text information is immediately subjected to a translation process.
[0828] Step 5:
[0829] The server translates audio text data and in-video text information into a second language specified by the user. This translation is optimized for real-time operation.
[0830] Step 6:
[0831] The server converts the translated text information into speech using speech synthesis technology and places it as subtitles on the video. This synthesized speech and text are synchronized with the original video.
[0832] Step 7:
[0833] The user's device captures their facial expressions with a camera, and an emotion engine analyzes the user's emotions in real time based on this data. The emotion engine then sends the analysis results to the server.
[0834] Step 8:
[0835] Based on data received from the emotion engine, the server adaptively adjusts the translated content according to the user's emotions. For example, if the user is confused, the server will deliver content that emphasizes explanatory parts or plays audio at a slower pace.
[0836] Step 9:
[0837] The server streams the adjusted and translated video to the user's device in real time.
[0838] Step 10:
[0839] The user's device receives translated video streamed from the server and plays it back while synchronizing the video and audio. This allows the user to instantly view content adjusted to their own language.
[0840] (Example 2)
[0841] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0842] It is necessary to remove language barriers in international lectures and classes and provide support to help participants gain a better understanding. Furthermore, there is a need to improve learning effectiveness by providing flexible content that responds to participants' emotional states.
[0843] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0844] In this invention, the server includes means for converting audio data into text data in real time using speech recognition technology and translating the text data from a first language to a second language; means for extracting character information from video data using optical character recognition technology and translating the extracted character information from a first language to a second language; means for converting the translation result into audio data using speech synthesis technology and constructing a generated video by overlaying the audio data and translated text onto video data; means for adjusting the display timing and method of the generated video based on emotion data obtained from a device that analyzes the user's emotions; and means for delivering the generated video to an information terminal in real time. As a result, the user can receive translated content adapted to their emotional state in real time.
[0845] "Speech recognition technology" is a technology that converts human speech into text data using a computer system.
[0846] "Optical character recognition technology" is a technology that extracts character information from image or video data.
[0847] "Translation means" refers to a process or device for converting text data into another language.
[0848] "Speech synthesis technology" is a technology that converts text data into speech data and outputs it as natural-sounding speech.
[0849] A "generative AI model" is an artificial intelligence system that performs natural language processing to generate and translate text data.
[0850] A "sentiment analysis device" is a device that analyzes a user's emotions in real time and acquires that data.
[0851] "Video data" refers to digital data that contains visual information.
[0852] An "information terminal" is an electronic device that transmits, receives, and processes data via a network.
[0853] This invention relates to a system that translates the audio and video content of lectures and classes into other languages in real time and also has the function of recognizing the user's emotions. This system consists of a server, information terminals, an emotion analysis device, and a network connecting them.
[0854] The user first uses an information terminal to select the content they want to view and specify the desired translation language. This request is sent to the server. The server uses speech recognition technology (e.g., commercial speech recognition software) to convert the acquired speech data into text data.
[0855] Next, a generative AI model (for example, a general-purpose generative AI engine) is used to translate the text data into the specified language. A prompt such as "Translate Japanese into English" is used.
[0856] For video data, the server extracts text information using optical character recognition technology (e.g., common OCR software) and translates this information. The translation result is converted into speech using speech synthesis technology (e.g., standard text-to-speech software) and incorporated into the video as an overlay.
[0857] For sentiment analysis, the camera and microphone built into the user's information terminal are used. The data acquired is analyzed in real time by a sentiment analysis device (e.g., a standard sentiment recognition API), and the sentiment information is sent to a server. Based on the sentiment information, the server can dynamically adjust how the generated video is displayed.
[0858] A concrete example is a user who wants to understand a Spanish physics lecture in English. The user selects the relevant Spanish lecture from their device and specifies English as the translation language. The audio and video are recognized and translated, and the audio and text are provided to the user in English in a synchronized format. Furthermore, the server can appropriately adjust the playback speed and level of detail according to the user's level of difficulty.
[0859] In this way, this system enables users to understand educational content in other languages in a way that is optimized for their own emotions.
[0860] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0861] Step 1:
[0862] The user selects the lecture or class content they wish to view using their information terminal and specifies their desired translation language. A request is generated containing the user's selected video data and desired language information. This request is sent to the server, where processing begins.
[0863] Step 2:
[0864] The server retrieves a specified video stream over the network. It uses video data based on user requests as input. The output is generated by converting audio data into text data using speech recognition technology. Specifically, the speech recognition engine analyzes the audio waveform and outputs it as text in the corresponding language.
[0865] Step 3:
[0866] The server uses a generative AI model to translate the text data obtained in step 2 into the specified language. The recognized original text data and a prompt message are used as input. A prompt message such as "Please translate Japanese into English" is used. The output is the translated text data.
[0867] Step 4:
[0868] The server extracts text information from video data using optical character recognition (OCR) technology. The input is video data, and the output is the extracted text information. The OCR engine analyzes the characters in the image and outputs them as corresponding text.
[0869] Step 5:
[0870] The character information extracted in Step 4 is similarly translated into the specified language using the generative AI model. The recognized character information and prompt text are used as input, and the output is the translated character information.
[0871] Step 6:
[0872] The server uses speech synthesis technology to convert translated text data into new speech data. The input is translated text data, and the output is speech data in the specified language. The speech synthesis engine analyzes the text and outputs it as synthesized speech.
[0873] Step 7:
[0874] The server overlays the generated audio data and translated text onto the video data to construct the generated video. The input consists of the generated audio data, translated text data, and original video data, and the output is the generated video in which they are integrated.
[0875] Step 8:
[0876] The emotion analysis device uses data obtained from the camera and microphone built into the device to detect the user's emotions and sends it to the server as temporary input data. This data indicates the emotional state, and a specific emotional pattern is obtained as output.
[0877] Step 9:
[0878] The server adjusts how the generated video is displayed based on the user's emotion data. The input is emotion data and generated video data, and the output is the generated video with the adjusted playback parameters.
[0879] Step 10:
[0880] The server delivers the adjusted, generated video to the information terminal in real time. Users view the content played on their terminal, receiving translations and emotionally-responsive learning support. The final output is the content played back in real time.
[0881] (Application Example 2)
[0882] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0883] In an increasingly globalized world, understanding educational and informational content provided in different languages remains challenging, transcending language barriers. Furthermore, the lack of content adjustments tailored to the audience's emotions and level of understanding leads to decreased learning efficiency. Therefore, a system is needed that provides real-time multilingual translation and customized information based on the audience's emotions.
[0884] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0885] In this invention, the server includes means for processing audio data in real time and translating it from a first language to a second language, means for extracting text information from video data and translating it from a first language to a second language, and means for acquiring the user's emotional state and adjusting the display method of the generated video based on the emotion. This enables the display of multilingual educational content in a customized manner according to the viewer's emotions, resulting in a more effective learning experience.
[0886] "Means for processing audio data in real time and translating from a first language to a second language" refers to a device or function that performs the process of instantly analyzing an audio signal, converting it into text, and then translating that text into a specified different language.
[0887] "Means for extracting text information from video data and translating it from a first language to a second language" refers to a device or function that identifies and acquires characters within a video and processes those characters to represent them in a different language.
[0888] "Means for synchronizing translation results with video data and constructing generated video including translated audio and text" refers to a device or function that performs the process of combining audio and text translated into different languages with the original video to create integrated video content.
[0889] "Means for delivering generated video to terminals in real time" refers to a network system and its control process that transfers the created translated video to the viewer's device without delay.
[0890] "Means for acquiring the user's emotional state and adjusting the display method of generated video based on that emotion" refers to a device or function that analyzes the user's facial expressions and behavioral data to recognize emotions and automatically changes the video playback speed, display content, etc., according to the analysis results.
[0891] The system for realizing this invention is composed of an integrated set of technologies, primarily consisting of servers, user terminals, and various software technologies.
[0892] First, the server receives a viewing request from the user. This request includes the content to be viewed and the desired translation language. The server acquires the audio data and converts it to text using speech recognition technology. This process primarily uses Python and speech recognition libraries.
[0893] Furthermore, the server uses optical character recognition (OCR) technology to extract text information from the video data. Simultaneously, a generative AI model is used to translate this text into the language specified by the user. At this stage, the OpenAI API is often used.
[0894] The server uses the user's device's camera to analyze facial expressions and behavior in real time in order to capture the user's emotional state. Machine learning libraries such as TensorFlow are used for emotion analysis.
[0895] The converted audio and text are synchronized with the video data as translation results. Speech synthesis technology is used here, and the resulting video is displayed in a way that adjusts based on the user's emotional state. For example, if the user is determined to be confused, the playback speed can be adjusted, or additional detailed explanations can be displayed. Finally, the adjusted video and audio are delivered to the device.
[0896] For example, if a Japanese user watching a technical lecture in English uses this system, and it determines that technical terms are difficult to understand, the system will display a detailed explanation of those terms in Japanese and play the audio commentary at a slower pace. The following is an example of a prompt to the generative AI model:
[0897] "Translate the following educational content from English to Japanese, emphasizing clarity for a high school student: English lesson content"
[0898] This makes it possible to provide a high-quality learning experience that transcends language and cultural barriers.
[0899] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0900] Step 1:
[0901] The server receives viewing requests from users. The input includes the ID of the content they want to view and the desired translation language. Based on this information, the server retrieves the corresponding video and audio data.
[0902] Step 2:
[0903] The server sends the acquired audio data to the speech recognition engine, which converts it into text data. The input is an audio signal, and the output is the recognized text. This process is performed by a Python speech recognition library.
[0904] Step 3:
[0905] The server extracts text information from video data using optical character recognition (OCR) technology. The input is video frames, and the output is the string data contained within the video. An OCR library is used in this step.
[0906] Step 4:
[0907] The server sends text data to a generation AI model for translation into the desired language. The input is text data in the original language, and the output is text data translated into the specified language. This translation process uses the OpenAI API.
[0908] Step 5:
[0909] The user's device captures facial expressions using its built-in camera and sends them to an emotion recognition engine. The input is facial image data, and the output is data indicating the user's emotional state. Libraries such as TensorFlow are used for emotion analysis.
[0910] Step 6:
[0911] The server synchronizes the translation results with the video data based on sentiment data and adjusts the display accordingly. The input consists of the translation results and sentiment data, and the output is a video with adjusted display settings. Specifically, this includes changes to font size and playback speed.
[0912] Step 7:
[0913] The server delivers generated, translated, and emotion-sensitive video to the terminal. The input is the adjusted video, and the output is this video displayed in real time on the user's terminal. This utilizes streaming technology over a network.
[0914] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0915] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0916] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0917] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0918] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0919] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0920] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0921] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0922] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0923] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0924] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0925] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0926] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0927] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0928] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0929] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0930] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0931] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0932] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0933] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0934] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0935] The following is further disclosed regarding the embodiments described above.
[0936] (Claim 1)
[0937] A means of processing audio data in real time and translating it from one language to another,
[0938] A means of extracting text information from video data and translating it from the first language to the second language,
[0939] A means for synchronizing the translation results with video data and constructing a generated video that includes translated audio and text,
[0940] A means of delivering the generated video to the terminal in real time,
[0941] A system that includes this.
[0942] (Claim 2)
[0943] The system according to claim 1, comprising means for converting speech data into text using speech recognition technology.
[0944] (Claim 3)
[0945] The system according to claim 1, comprising means for extracting text information from video data using optical character recognition technology.
[0946] "Example 1"
[0947] (Claim 1)
[0948] A means for receiving requests to view lectures or classes from user terminals,
[0949] A means of acquiring video via a network based on a request and converting the audio track into text using speech recognition technology,
[0950] A means for extracting text information from video using optical character recognition technology,
[0951] A method for translating transcribed audio data and extracted text information in real time using multilingual translation technology,
[0952] A means of converting translated audio data into an audio file using speech synthesis technology,
[0953] A means of integrating translated audio and text into the original video data to generate video,
[0954] A means of delivering the generated video to the user's terminal in real time,
[0955] A system that includes this.
[0956] (Claim 2)
[0957] The system according to claim 1, comprising means for converting speech data into text using speech recognition technology.
[0958] (Claim 3)
[0959] The system according to claim 1, comprising means for extracting character information from video data using optical character recognition technology.
[0960] "Application Example 1"
[0961] (Claim 1)
[0962] A means of processing audio data in real time and translating it from one language to another,
[0963] A means of extracting text information from video data and translating it from the first language to the second language,
[0964] A means for synchronizing the translation result with video data and constructing a generated video that includes the translated audio and text,
[0965] A communication method that delivers the generated video to the terminal in real time and allows the user to specify a selective viewing language,
[0966] A means of providing educational information in the language of the country in the selected language and displaying it via a visual device,
[0967] A system that includes this.
[0968] (Claim 2)
[0969] The system according to claim 1, comprising means for converting speech data into text using speech recognition technology.
[0970] (Claim 3)
[0971] The system according to claim 1, comprising means for extracting text information from video data using optical character recognition technology.
[0972] "Example 2 of combining an emotion engine"
[0973] (Claim 1)
[0974] A means of converting speech data into text data in real time using speech recognition technology, and translating that text data from a first language to a second language,
[0975] A means for extracting text information from video data using optical character recognition technology and translating the extracted text information from a first language to a second language,
[0976] A means for constructing a generated video by converting the translation result into audio data using speech synthesis technology, and overlaying that audio data and the translated text onto video data,
[0977] A means for adjusting the display timing and method of the generated video based on emotional data obtained from a device that analyzes user emotions,
[0978] A means of distributing the generated video to information terminals in real time,
[0979] A system that includes this.
[0980] (Claim 2)
[0981] The system according to claim 1, comprising means for translating text data into a specified language using a generative AI model.
[0982] (Claim 3)
[0983] The system according to claim 1, comprising means for detecting the user's emotions using an emotion analysis device and changing the display of the video based on those emotions.
[0984] "Application example 2 when combining with an emotional engine"
[0985] (Claim 1)
[0986] A means of processing audio data in real time and translating it from one language to another,
[0987] A means of extracting text information from video data and translating it from the first language to the second language,
[0988] A means for synchronizing the translation results with video data and constructing a generated video that includes translated audio and text,
[0989] A means of delivering the generated video to the terminal in real time,
[0990] A means for acquiring the user's emotional state and adjusting the display method of the generated video based on that emotion,
[0991] A system that includes this.
[0992] (Claim 2)
[0993] The system according to claim 1, comprising means for converting speech data into text using speech recognition technology.
[0994] (Claim 3)
[0995] The system according to claim 1, comprising means for extracting text information from video data using optical character recognition technology. [Explanation of symbols]
[0996] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of processing audio data in real time and translating it from one language to another, A means of extracting text information from video data and translating it from the first language to the second language, A means for synchronizing the translation results with video data and constructing a generated video that includes translated audio and text, A means of delivering the generated video to the terminal in real time, A system that includes this.
2. The system according to claim 1, comprising means for converting speech data into text using speech recognition technology.
3. The system according to claim 1, comprising means for extracting text information from video data using optical character recognition technology.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A