system
The system addresses the limitations of conventional subtitle generation by extracting audio, identifying speakers, and outputting subtitles in editable formats, enhancing visibility and editing efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Conventional subtitle generation methods fail to organize key points, distinguish speakers, and provide editable subtitle files, leading to low visibility and difficulty in editing.
A system that extracts audio from video data, converts it into text, identifies speakers, and generates subtitles with different fonts and colors, outputting in an editable project file format.
Enables efficient, highly legible subtitle generation that captures key points and allows further editing, improving the visibility and editing efficiency of subtitles.
Smart Images

Figure 2026070186000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, with the increase in video distribution, the demand for subtitles has been increasing. However, with conventional subtitle generation means, the key points are not organized just by transcribing the audio as it is, resulting in subtitles that are difficult for viewers to understand. Also, since the font and color of the subtitles are uniform, it is difficult to distinguish speakers, and there is a problem of low visibility. Furthermore, the generated subtitle file is only output in text format, causing constraints when freely editing with video editing software. The present invention aims to solve these problems and provide an efficient and highly visible subtitle generation method.
Means for Solving the Problems
[0005] This invention extracts audio from input video data and converts it into text data using speech recognition technology. Furthermore, it provides a means for extracting key points from the text data to generate subtitles, automatically identifying speakers, and displaying subtitles with different fonts and colors for each speaker. In addition, it facilitates final editing work by outputting the subtitle data in an editable project file format. This enables the automatic generation of concise, highly legible subtitles that capture the key points, and allows for further editing using other video editing software.
[0006] "Video data" refers to multimedia content in a format that combines video and audio.
[0007] "Audio extraction" is the process of separating only the audio portion from video data.
[0008] "Speech recognition technology" is a technology that analyzes speech data and converts it into text data.
[0009] "Text data" refers to a form of digital information that uses characters.
[0010] "Key point extraction" is the process of selecting only the important parts from a large amount of information.
[0011] Subtitles are textual information displayed over a video, providing supplementary explanations of the video's content.
[0012] "Speaker identification" is the process of identifying the speaker from audio data.
[0013] A "font" refers to the design of the shape and style of letters.
[0014] "Color" is a characteristic that is expressed by a combination of the three primary colors of light that are perceived visually.
[0015] "Project file format" refers to a method of saving data in a specific format usable by video editing software.
Brief Description of the Drawings
[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Modes for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] The system according to the present invention is for automatically generating effective subtitles for videos and providing them to users. This system consists of a server, a terminal, and a user interface.
[0038] First, the user uploads the video data for which they want subtitles to be generated to the system using their own device. The server extracts the audio from the received video data and uses speech recognition technology to convert that audio data into text data.
[0039] For example, in the case of a video recording of a lecture, the speaker's words are transcribed sequentially. The server then analyzes the generated text data and extracts the key points using natural language processing. This generates subtitles that highlight the important points of the lecture.
[0040] Furthermore, the server has a function to automatically identify speakers from audio data, and it is configured to display subtitles in different fonts and colors for each speaker. For example, in a video in a dialogue format, different speakers can be visually distinguished by applying blue and green fonts to each.
[0041] The generated subtitles are output in a project file format that can be edited with video editing software such as Final Cut Pro and Premiere Pro, according to the user's needs, allowing users to easily perform video editing tasks. Users can save the completed project file to their device and make any necessary adjustments to their final video work.
[0042] This system generates highly legible subtitles in a short time through an automated process, allowing users to achieve a professional finish without requiring specialized skills. Thus, the present invention aims to improve the efficiency of subtitle generation and editing in video production.
[0043] The following describes the processing flow.
[0044] Step 1:
[0045] The user uploads video data requiring subtitles to the server using their device. During this process, the user specifies the video file by dragging and dropping it using the system interface.
[0046] Step 2:
[0047] The server extracts the audio track from the received video data. This process separates the audio portion of the video file and saves it as an audio file.
[0048] Step 3:
[0049] The server analyzes the extracted audio using speech recognition technology and converts the audio data into text data. This conversion involves transcribing the content of the audio utterance by utterance, creating text from what was said.
[0050] Step 4:
[0051] The server analyzes the generated text data using natural language processing techniques and extracts the key points of the text to produce concise and clear subtitles. This creates subtitles that highlight particularly important information in the audio.
[0052] Step 5:
[0053] The server identifies the speaker from the audio data and sets different fonts and colors for each speaker. This identification makes it possible to visually distinguish between statements made by different speakers.
[0054] Step 6:
[0055] The server outputs the generated subtitle data as a project file format editable in Final Cut Pro or Premiere Pro. This project file contains the subtitle information and its styling.
[0056] Step 7:
[0057] Users download project files from the server and make final adjustments to subtitles using video editing software on their own devices. Within the editing environment, they can freely change the position and style of the subtitles and incorporate them into the final video.
[0058] (Example 1)
[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0060] In current video production, generating visually clear subtitles requires considerable effort and specialized skills. Furthermore, differentiating subtitle display for each speaker and creating subtitles that capture the key points also takes time and effort, necessitating an efficient process.
[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0062] In this invention, the server includes means for extracting audio information from video information, means for converting the audio information into text information using a speech recognition method, and means for extracting important information from the text information and generating display information that suppresses the important information. This makes it possible to quickly and effectively generate highly legible subtitles without requiring specialized knowledge.
[0063] "Video information" refers to information that includes video data and related audio data.
[0064] "Audio information" refers to audio data extracted from video information.
[0065] "Speech recognition method" is a technology that analyzes speech information and converts it into text information.
[0066] "Textual information" refers to text data obtained through speech recognition methods.
[0067] "Important information" refers to information that contains the main content or key points extracted from textual information.
[0068] "Display information" refers to subtitles and text display data generated based on important information.
[0069] "Speakers" refer to the individual individuals who are speaking within the audio information.
[0070] "Display format" refers to styles and formats used to visually change the appearance of text and graphics.
[0071] "Project information format" refers to a file format that saves display information in a format that allows for video editing.
[0072] "Rules" refer to established rules and standards for setting display formats based on speaker identification.
[0073] "Opinions" refers to feedback and evaluations from users.
[0074] This invention relates to a system for automatically generating highly legible subtitles from video information, and its main components are a user, a terminal, and a server. First, the user uploads video information requiring subtitles to the server using their own terminal. This includes file selection and upload operations via a web browser or a specific application.
[0075] The server analyzes the received video information and extracts the audio information. Audio processing libraries such as FFmpeg are used for this audio extraction. Next, the server applies a speech recognition method to convert the audio information into text information. Here, speech recognition technologies such as Google's Speech-to-Text API are used.
[0076] The converted text information is processed by a server, and important information is extracted by a natural language processing engine (e.g., spaCy). This generates display information containing the key points. The generated display information is then adjusted by speaker identification technology to output in a different display format for each speaker.
[0077] As a final result, the displayed information is output in project information format and can be downloaded to the user's device. Since this format is compatible with video editing software such as Final Cut Pro and Premiere Pro, users can further edit subtitles using this format.
[0078] As a concrete example, suppose a user uploads an educational video in which an instructor explains different topics. In this case, displaying the key points of each topic immediately after the start of each topic will make it easier for viewers to understand the content.
[0079] An example of a prompt to input into the generation AI model would be, "Generate the most relevant information for this educational video and visually highlight the key points of each section." This system allows users to quickly generate visually effective subtitles without requiring specialized knowledge or skills.
[0080] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0081] Step 1:
[0082] Users upload video information to the server using their own devices to generate subtitles. Specifically, users select video files and upload them using a dedicated application or web browser interface. The input is the video file, and the output is that file stored on the server.
[0083] Step 2:
[0084] The server extracts audio information from the uploaded video data. This extraction uses an audio processing library such as FFmpeg. This library extracts the audio track from the video file and writes it as an audio file. The input is a video file, and the output is a separated audio file.
[0085] Step 3:
[0086] The server uses speech recognition to convert extracted audio information into text. This process utilizes the Google Speech-to-Text API, sending audio data and receiving corresponding text data. The input is an audio file, and the output is text data.
[0087] Step 4:
[0088] The server applies natural language processing to the generated text information to extract important information. For example, it uses a natural language processing library such as spaCy to perform grammatical analysis and key phrase extraction. The input is text data, and the output is important information with key points highlighted.
[0089] Step 5:
[0090] The server uses speaker identification technology to adjust the display format of important information for each speaker. For example, it performs speech feature analysis and applies different fonts and colors based on each speaker's speech pattern. The input is important information, and the output is data displayed in a format appropriate for the speaker.
[0091] Step 6:
[0092] The server outputs the display information in project information format, making it available for download on the user's device. The project file is generated in a format that can be read by video editing software. The input is display data adjusted for each speaker, and the output is a project information file.
[0093] Users can download the project information file and perform further editing on their own devices.
[0094] (Application Example 1)
[0095] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0096] In modern content distribution services, viewers watch videos in a variety of situations, and there is a growing need for automatic subtitle generation to aid visual comprehension, especially for videos with difficult-to-understand audio or in different languages. Therefore, there is a demand for real-time, high-quality, and visually clear subtitle generation to improve the viewing experience.
[0097] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0098] In this invention, the server includes means for extracting audio from input video data, means for converting audio into text data using speech recognition technology, and control means for generating and visually presenting display information in real time. This allows viewers to visually obtain high-quality subtitles that are automatically generated in real time, making it easier to understand the content they are watching.
[0099] "Video data" refers to digital data that includes visual information, including dynamic screen information such as videos.
[0100] "Speech recognition technology" is a technology that analyzes collected speech data and converts it into text information based on that analysis.
[0101] "Text data" refers to text information converted by speech recognition technology, and is data expressed in a form that is readable by humans.
[0102] "Display information" refers to data presented visually, particularly subtitles and supplementary information displayed within a video.
[0103] "Speaker identification" is the process of identifying and distinguishing different speakers, and it is performed using the characteristics of voices and sounds within the video.
[0104] "Project data" refers to a format of editable data used in video editing, and includes a set of data files containing subtitles and other video editing elements.
[0105] "Control means" refers to mechanisms and programs for managing and operating various processes in an image processing system, enabling real-time processing.
[0106] "Real-time" refers to a temporal concept where instructed actions or processes are carried out immediately without delay, signifying an immediate response.
[0107] This invention provides a real-time subtitle generation system for a video distribution platform. This system consists of a server, a terminal, and a user interface.
[0108] The server first receives video data for the user to view. When extracting audio from the video data, it uses speech recognition technology to convert the audio into text data. For example, the Python `speech_recognition` library can be used to perform this audio conversion efficiently.
[0109] The converted text data is analyzed on the server and visually presented as real-time display information. This includes compositing subtitles onto the video using video editing software such as the moviepy library. Each speaker's statements are displayed in a different color and font, making it easy for viewers to distinguish between speakers.
[0110] The device instantly visualizes the generated display information on the user's device, allowing the user to view high-quality subtitles in real time while watching. This makes it easier to understand the content being viewed.
[0111] For example, when a user is watching a cooking video, automatically generated subtitles are displayed by the system, allowing viewers to more clearly understand the steps of the recipe. To further enhance user convenience, the subtitle format and display style can be customized through the device settings.
[0112] An example of a prompt message is: "Generate subtitles for a cooking video. Create subtitles that clearly indicate ingredients and steps, and visually distinguish between different speakers (chef and narrator)." This prompt allows users to easily generate display information tailored to their specific needs.
[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0114] Step 1:
[0115] The server receives video data being streamed from the user and extracts audio from that data. This process uses video data as input and outputs an audio stream. The extraction of audio data involves using the necessary APIs and libraries to obtain the audio file.
[0116] Step 2:
[0117] The server uses speech recognition technology to convert the extracted audio stream into text data. This process takes the audio stream as input and obtains text data as output. The speech_recognition library is used to perform speech analysis and convert the spoken portion into text.
[0118] Step 3:
[0119] The server analyzes the converted text data and generates subtitles as display information. Here, text data is used as input, and subtitles in a visually displayable format are output. Subtitle generation uses natural language processing techniques to extract important points and determine the display format based on them.
[0120] Step 4:
[0121] The terminal receives the generated subtitle data and displays it to the user in real time. The generated subtitle data is used as input, and the real-time display is output on the user's device. Libraries such as moviepy are used to perform the specific action of overlaying the subtitles onto the display screen.
[0122] Step 5:
[0123] Users visually review the subtitles presented to understand the content they are watching. They directly input the generated display information and receive real-time improvements to their viewing experience as output. This process involves using the displayed information to understand and process the viewing experience.
[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0125] The system according to the present invention not only generates effective subtitles for videos, but also has the function of recognizing the user's emotions and reflecting them in the subtitles. This system mainly consists of a server, a terminal, and a user interface, and is realized through the following process.
[0126] The user uploads video data to the system using their own device. The server separates the audio from the uploaded video and converts it into text data using speech recognition technology. At this point, the text data is processed to generate concise subtitles.
[0127] Furthermore, this invention uses an emotion engine to recognize the user's emotions. This emotion recognition is performed by analyzing the user's facial expressions and voice tone through devices such as a camera and microphone connected to the terminal. For example, if the user shows a surprised expression, the emotion engine recognizes this and makes changes such as emphasizing the subtitle display style.
[0128] Based on the emotions it recognizes, the server dynamically adjusts the subtitle styling (e.g., changing colors and fonts) and display time, and further improves the accuracy of key point extraction for specific scenes. The subtitles generated in this way are output as a project file that can be freely edited with video editing software such as Final Cut Pro and Premiere Pro, and are provided to the user.
[0129] Users can import these project files into their video editing software and adjust them to best suit the context of their video work. In this way, the system of the present invention makes it possible to make the video viewing experience more interactive and engaging by utilizing emotional information.
[0130] The following describes the processing flow.
[0131] Step 1:
[0132] The user accesses the system interface from their terminal, selects the video data for which they want to generate subtitles, and uploads it to the server. The user interface provides a guide to simplify the process.
[0133] Step 2:
[0134] The server extracts audio from the uploaded video data. This process separates the audio track from the video file and converts it into a format suitable for analysis.
[0135] Step 3:
[0136] The server uses speech recognition technology to convert the extracted speech into text data. The spoken content is sequentially converted into text information through speech analysis.
[0137] Step 4:
[0138] The server extracts key points from the generated text data and produces clear subtitles. Natural language processing technology identifies relevant information and organizes it into a concise and easy-to-read format.
[0139] Step 5:
[0140] The device's built-in camera and microphone detect the user's facial expressions and voice tone, which are then analyzed by an emotion engine. For example, if the user smiles, that emotion is recognized as "joy."
[0141] Step 6:
[0142] The server adjusts subtitles based on the recognized user's emotions. By softening the color of the text or changing the font style, it provides readability that matches the user's emotions.
[0143] Step 7:
[0144] The server converts the final subtitle data into an editable project file format and provides the user with a download link. Users can download it to their own devices and make further adjustments using video editing software.
[0145] (Example 2)
[0146] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0147] When providing subtitles for video content, it is difficult to generate dynamic subtitle styles that reflect the viewer's emotional state, resulting in a challenge in adequately improving the immersion and quality of the viewing experience. Furthermore, there are limited ways to effectively incorporate user feedback.
[0148] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0149] In this invention, the server includes means for extracting audio from input video information, means for converting audio into text information using a speech recognition method, means for extracting key points from the text information and generating subtitles that highlight the key points, means for recognizing the user's emotions and dynamically adjusting the subtitle style based on the recognized emotions, and means for outputting the subtitle data in an editable format. This makes it possible to provide subtitles that reflect the user's emotional state, improve the quality of the viewing experience, and incorporate user feedback into the subtitle generation process.
[0150] "Input video information" refers to digital data, including visual and audio data, that the user provides to the system.
[0151] "Means for extracting audio" refers to a technology or device for separating and acquiring only audio data from video information.
[0152] "Speech recognition techniques" are algorithms and technologies used to convert speech data into strings of characters that a computer can understand.
[0153] "Textual information" refers to text data converted from audio data using speech recognition techniques, and is the information that forms the basis for subtitle generation.
[0154] "Means for extracting key points and generating concise subtitles" refers to algorithms or devices that analyze textual information, identify important content, and create subtitles based on that content.
[0155] "A means of recognizing user emotions and dynamically adjusting subtitle styles based on those emotions" refers to a technology that analyzes user emotions and changes the subtitle display format in real time based on the results.
[0156] "Means of outputting subtitle data in an editable format" refers to technologies and systems that provide generated subtitles in a file format that can be further edited using video editing software, etc.
[0157] This invention provides a system for generating emotion-based interactive subtitles for video content. This system primarily consists of a server, terminals, and a user interface.
[0158] Users can upload video data they want to view or edit to the system using their own devices. The server extracts audio from the received video data using multimedia tools and converts it into text using speech recognition technology. Common speech recognition APIs are used in this process.
[0159] Next, the server extracts key points from the text information and generates subtitles based on that content. A natural language processing model is used in this process. The generated subtitles are then styled in real time on the server according to the user's emotions, which are analyzed by an emotion engine via the camera and microphone on the user's device. For example, if the user smiles or shows a surprised expression, the color and font of the subtitles will change accordingly.
[0160] Furthermore, the server can output the generated subtitles as a project file that can be edited using the user's editing software. This allows users to freely add to and edit the provided subtitles to use them in a way that is more appropriate to the context of the video.
[0161] As a concrete example, imagine a scenario where a user uploads a travel video, and a system that detects the user's smile generates a caption saying "The joy of sharing this moment!" and changes the layout to a brighter color scheme.
[0162] An example of a prompt is, "Generate emotion-recognition-based subtitles for a video of our summer vacation together as a family." Using this prompt, users can naturally create subtitles that reflect their own emotions in the video.
[0163] Thus, the present invention provides subtitles that reflect the user's emotions, making the video viewing experience more interactive and engaging.
[0164] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0165] Step 1:
[0166] The user uploads video data to the system using their own device. The input for this step is video data, and the output is the transmission of data to the server. The user selects a file on the operation screen and sends that file to the system.
[0167] Step 2:
[0168] The server extracts audio from the uploaded video data. The input for this step is the video data sent by the user, and the output is the extracted audio data. The server uses multimedia processing tools to analyze the video data and perform the specific operation of separating the audio track.
[0169] Step 3:
[0170] The server converts the extracted audio data into text information using speech recognition technology. The input for this step is audio data, and the output is the converted text information. The server calls a speech recognition API to perform specific calculations that sequentially process the audio data and convert it into text.
[0171] Step 4:
[0172] The server extracts key points from text information to generate subtitles. The input for this step is text information, and the output is summarized subtitle data. The server utilizes natural language processing techniques to analyze the text, identify important key phrases, and perform the specific processing required to form subtitles.
[0173] Step 5:
[0174] The device recognizes the user's emotions and sends that data to the server. The input for this step is the user's facial expressions and voice tone collected from the device's camera and microphone, and the output is the recognized emotion data. The device uses emotion analysis software to perform specific calculations to determine the user's emotional state in real time.
[0175] Step 6:
[0176] The server dynamically adjusts the subtitle style based on the received sentiment data. The input for this step is the sentiment data and the generated subtitle data, and the output is the final, styled subtitle data. The server refers to style rules and performs specific operations to change the font, color, and display time of the subtitles.
[0177] Step 7:
[0178] The server outputs subtitle data in an editable format and provides it to the user. The input for this step is styled subtitle data, and the output is a project file usable with video editing software. The server uses a file format conversion tool to save the data in the appropriate format, making it available for user download.
[0179] (Application Example 2)
[0180] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0181] In recent years, as video content viewing has diversified, there has been a growing demand for interactive subtitles that respond to viewers' emotional states. However, conventional video subtitling systems have lacked the ability to reflect individual viewers' emotions, making it difficult to provide a deeper viewing experience. Furthermore, there has been a lack of subtitle editing technology that can reflect viewers' emotions in real time. Solving these problems is a challenge.
[0182] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0183] In this invention, the server includes means for extracting audio from input video data, means for converting audio into text data, and means for sensing emotions using an input device to analyze emotional states. This enables a more interactive and personalized viewing experience by dynamically adjusting subtitle styling and display time based on the viewer's emotional information.
[0184] An "input device" is a device used to acquire data such as audio and video and convert it into a format that can be used within the system.
[0185] "Speech recognition technology" is a technology that analyzes speech data and converts it into corresponding text data.
[0186] "Text data" refers to character information extracted from speech using speech recognition technology.
[0187] "Styling" refers to adjusting display attributes such as the color, font, and size of subtitles.
[0188] "Display characteristics" refer to visual features such as the color scheme and fonts of subtitles and interfaces.
[0189] "Emotion" refers to the psychological state, such as feelings and moods, that arise in the user's mind.
[0190] "Data format" refers to a standardized structure or format used when storing or processing digital information.
[0191] "Interaction" refers to the two-way interaction and exchange of information that takes place between a system and a user.
[0192] "Real-time" refers to a time concept where information is processed instantly, enabling immediate responses.
[0193] To realize this invention, three main components are necessary: a server, a terminal, and a user. The server extracts audio from video data and converts that audio into text data using speech recognition technology. Examples of speech recognition technologies used in this process include the Google Cloud Speech-to-Text API. Key points are extracted from the generated text data, and subtitles are dynamically generated.
[0194] The device plays a role in sensing the user's emotional state in real time. For this purpose, the device uses a built-in camera and microphone to analyze the user's facial expressions and tone of voice. The software used for analyzing emotional states utilizes the Azure® Emotion Recognition API. This allows for the collection of real-time emotional data from the user, and the dynamic adjustment of subtitle styling and display time.
[0195] Users operate their own devices and upload video data to the server. The processed subtitles are provided in a data format that can be edited with video editing software such as Adobe Premiere Pro. In this way, the video viewing experience can be made more interactive and personalized.
[0196] For example, if a user is moved during the climax scene of a movie, the system can change the subtitle color to a softer shade that matches their emotion and even suggest an interaction such as, "Would you like to share this feeling with a friend?"
[0197] An example of a prompt sentence to input into a generative AI model is, "Please suggest a suitable subtitle style for this scene in this movie."
[0198] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0199] Step 1:
[0200] The user uploads video data from their device to the server. The uploaded video data is passed to the server as input. The server receives this data and prepares to analyze the video file.
[0201] Step 2:
[0202] The server extracts audio from the video data. Using the extracted audio data as input, it generates an audio file. This process transforms the video's audio information into a separate file, preparing it for the subsequent speech recognition process.
[0203] Step 3:
[0204] The server uses the Google Cloud Speech-to-Text API to convert the extracted audio data into text data. This text data is output, providing information in written form. Speech recognition technology transforms spoken content into readable text data.
[0205] Step 4:
[0206] The server analyzes the generated text data and applies an algorithm to extract key points. It extracts key points from the input text data and uses them to generate concise subtitle data that highlights the essentials. This results in a concise subtitle containing only the most important information.
[0207] Step 5:
[0208] The device senses the user's emotional state in real time using its camera and microphone. This sensor data is then used to analyze the user's emotional state using the Azure Emotion Recognition API. The analysis results are output, and the emotional state is quantified.
[0209] Step 6:
[0210] The server dynamically adjusts subtitle styling and display time based on user emotion data. Emotion data is used as input, and the color, font, and timing of subtitle display are changed in real time. This creates subtitles that visually reflect emotions.
[0211] Step 7:
[0212] The server generates adjusted subtitle data and outputs it in a format editable with Adobe Premiere Pro and other software. The output project file format is then provided to the user, allowing them to edit the video according to their own needs.
[0213] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0214] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0215] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0216] [Second Embodiment]
[0217] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0218] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0219] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0220] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0221] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0222] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0223] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0224] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0225] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0226] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0227] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0228] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0229] The system according to the present invention is for automatically generating effective subtitles for videos and providing them to users. This system consists of a server, a terminal, and a user interface.
[0230] First, the user uploads the video data for which they want subtitles to be generated to the system using their own device. The server extracts the audio from the received video data and uses speech recognition technology to convert that audio data into text data.
[0231] For example, in the case of a video recording of a lecture, the speaker's words are transcribed sequentially. The server then analyzes the generated text data and extracts the key points using natural language processing. This generates subtitles that highlight the important points of the lecture.
[0232] Furthermore, the server has a function to automatically identify speakers from audio data, and it is configured to display subtitles in different fonts and colors for each speaker. For example, in a video in a dialogue format, different speakers can be visually distinguished by applying blue and green fonts to each.
[0233] The generated subtitles are output in a project file format that can be edited with video editing software such as Final Cut Pro and Premiere Pro, according to the user's needs, allowing users to easily perform video editing tasks. Users can save the completed project file to their device and make any necessary adjustments to their final video work.
[0234] This system generates highly legible subtitles in a short time through an automated process, allowing users to achieve a professional finish without requiring specialized skills. Thus, the present invention aims to improve the efficiency of subtitle generation and editing in video production.
[0235] The following describes the processing flow.
[0236] Step 1:
[0237] The user uploads video data requiring subtitles to the server using their device. During this process, the user specifies the video file by dragging and dropping it using the system interface.
[0238] Step 2:
[0239] The server extracts the audio track from the received video data. This process separates the audio portion of the video file and saves it as an audio file.
[0240] Step 3:
[0241] The server analyzes the extracted audio using speech recognition technology and converts the audio data into text data. This conversion involves transcribing the content of the audio utterance by utterance, creating text from what was said.
[0242] Step 4:
[0243] The server analyzes the generated text data using natural language processing techniques and extracts the key points of the text to produce concise and clear subtitles. This creates subtitles that highlight particularly important information in the audio.
[0244] Step 5:
[0245] The server identifies the speaker from the audio data and sets different fonts and colors for each speaker. This identification makes it possible to visually distinguish between statements made by different speakers.
[0246] Step 6:
[0247] The server outputs the generated subtitle data as a project file format editable in Final Cut Pro or Premiere Pro. This project file contains the subtitle information and its styling.
[0248] Step 7:
[0249] Users download project files from the server and make final adjustments to subtitles using video editing software on their own devices. Within the editing environment, they can freely change the position and style of the subtitles and incorporate them into the final video.
[0250] (Example 1)
[0251] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0252] In current video production, generating visually clear subtitles requires considerable effort and specialized skills. Furthermore, differentiating subtitle display for each speaker and creating subtitles that capture the key points also takes time and effort, necessitating an efficient process.
[0253] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0254] In this invention, the server includes means for extracting audio information from video information, means for converting the audio information into text information using a speech recognition method, and means for extracting important information from the text information and generating display information that suppresses the important information. This makes it possible to quickly and effectively generate highly legible subtitles without requiring specialized knowledge.
[0255] "Video information" refers to information that includes video data and related audio data.
[0256] "Audio information" refers to audio data extracted from video information.
[0257] "Speech recognition method" is a technology that analyzes speech information and converts it into text information.
[0258] "Textual information" refers to text data obtained through speech recognition methods.
[0259] "Important information" refers to information that contains the main content or key points extracted from textual information.
[0260] "Display information" refers to subtitles and text display data generated based on important information.
[0261] "Speakers" refer to the individual individuals who are speaking within the audio information.
[0262] "Display format" refers to styles and formats used to visually change the appearance of text and graphics.
[0263] "Project information format" refers to a file format that saves display information in a format that allows for video editing.
[0264] "Rules" refer to established rules and standards for setting display formats based on speaker identification.
[0265] "Opinions" refers to feedback and evaluations from users.
[0266] This invention relates to a system for automatically generating highly legible subtitles from video information, and its main components are a user, a terminal, and a server. First, the user uploads video information requiring subtitles to the server using their own terminal. This includes file selection and upload operations via a web browser or a specific application.
[0267] The server analyzes the received video information and extracts the audio information. Audio processing libraries such as FFmpeg are used for this audio extraction. Next, the server applies a speech recognition method to convert the audio information into text information. Here, speech recognition technologies such as the Google Speech-to-Text API are used.
[0268] The converted text information is processed by a server, and important information is extracted by a natural language processing engine (e.g., spaCy). This generates display information containing the key points. The generated display information is then adjusted by speaker identification technology to output in a different display format for each speaker.
[0269] As a final result, the displayed information is output in project information format and can be downloaded to the user's device. Since this format is compatible with video editing software such as Final Cut Pro and Premiere Pro, users can further edit subtitles using this format.
[0270] As a concrete example, suppose a user uploads an educational video in which an instructor explains different topics. In this case, displaying the key points of each topic immediately after the start of each topic will make it easier for viewers to understand the content.
[0271] An example of a prompt to input into the generation AI model would be, "Generate the most relevant information for this educational video and visually highlight the key points of each section." This system allows users to quickly generate visually effective subtitles without requiring specialized knowledge or skills.
[0272] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0273] Step 1:
[0274] Users upload video information to the server using their own devices to generate subtitles. Specifically, users select video files and upload them using a dedicated application or web browser interface. The input is the video file, and the output is that file stored on the server.
[0275] Step 2:
[0276] The server extracts audio information from the uploaded video data. This extraction uses an audio processing library such as FFmpeg. This library extracts the audio track from the video file and writes it as an audio file. The input is a video file, and the output is a separated audio file.
[0277] Step 3:
[0278] The server uses a speech recognition method to convert the extracted voice information into character information. For this process, the Google Speech-to-Text API is utilized to send the voice data and receive the corresponding character data. The input is a voice file, and the output is text data.
[0279] Step 4:
[0280] The server applies natural language processing to the generated character information to extract important information. For example, a natural language processing library like spaCy is used to perform syntactic analysis and key phrase extraction. The input is text data, and the output is important information with emphasized key points.
[0281] Step 5:
[0282] The server uses speaker identification technology to adjust the important information in different display formats for each speaker. For example, voice feature analysis is performed, and different fonts and colors are applied based on the vocal patterns of each speaker. The input is important information, and the output is data in a display format corresponding to the speaker.
[0283] Step 6:
[0284] The server outputs the display information in the form of project information, enabling it to be downloaded on the user's terminal. The project file is generated in a format that can be read by video editing software. The input is the display data adjusted for each speaker, and the output is a project information file.
[0285] The user can download this project information file and further perform editing work on their own terminal.
[0286] (Application Example 1)
[0287] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0288] In modern content distribution services, viewers watch videos in a variety of situations, and there is a growing need for automatic subtitle generation to aid visual comprehension, especially for videos with difficult-to-understand audio or in different languages. Therefore, there is a demand for real-time, high-quality, and visually clear subtitle generation to improve the viewing experience.
[0289] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0290] In this invention, the server includes means for extracting audio from input video data, means for converting audio into text data using speech recognition technology, and control means for generating and visually presenting display information in real time. This allows viewers to visually obtain high-quality subtitles that are automatically generated in real time, making it easier to understand the content they are watching.
[0291] "Video data" refers to digital data that includes visual information, including dynamic screen information such as videos.
[0292] "Speech recognition technology" is a technology that analyzes collected speech data and converts it into text information based on that analysis.
[0293] "Text data" refers to text information converted by speech recognition technology, and is data expressed in a form that is readable by humans.
[0294] "Display information" refers to data presented visually, particularly subtitles and supplementary information displayed within a video.
[0295] "Speaker identification" is the process of identifying and distinguishing different speakers, and it is performed using the characteristics of voices and sounds within the video.
[0296] "Project data" refers to a format of editable data used in video editing, and includes a set of data files containing subtitles and other video editing elements.
[0297] "Control means" refers to mechanisms and programs for managing and operating various processes in an image processing system, enabling real-time processing.
[0298] "Real-time" refers to a temporal concept where instructed actions or processes are carried out immediately without delay, signifying an immediate response.
[0299] This invention provides a real-time subtitle generation system for a video distribution platform. This system consists of a server, a terminal, and a user interface.
[0300] The server first receives video data for the user to view. When extracting audio from the video data, it uses speech recognition technology to convert the audio into text data. For example, the Python `speech_recognition` library can be used to perform this audio conversion efficiently.
[0301] The converted text data is analyzed on the server and visually presented as real-time display information. This includes compositing subtitles onto the video using video editing software such as the moviepy library. Each speaker's statements are displayed in a different color and font, making it easy for viewers to distinguish between speakers.
[0302] The device instantly visualizes the generated display information on the user's device, allowing the user to view high-quality subtitles in real time while watching. This makes it easier to understand the content being viewed.
[0303] As a specific example, when a user is watching a cooking video, subtitles automatically generated by the system are displayed, so that viewers can more clearly understand the steps of the recipe. To further enhance the convenience for the user, the format and display form of the subtitles can be customized through the settings of the terminal.
[0304] As an example of a prompt sentence, there is "Please generate subtitles for the cooking video. Create subtitles that clearly express the ingredients and steps and visually distinguish different speakers (the chef and the narrator).". With this prompt, the user can easily generate display information according to specific needs.
[0305] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0306] Step 1:
[0307] The server receives the video data being distributed from the user and extracts the audio from that data. In this process, video data is used as the input data and an audio stream is output. For the extraction of audio data, a process of obtaining an audio file is performed using the necessary APIs and libraries.
[0308] Step 2:
[0309] The server uses speech recognition technology to convert the extracted audio stream into character data. In this process, the audio stream is used as the input and character data is obtained as the output. The speech_recognition library is utilized to perform speech analysis and an operation to convert the spoken part into text is carried out.
[0310] Step 3:
[0311] The server analyzes the converted text data and generates subtitles as display information. Here, text data is used as input, and subtitles in a visually displayable format are output. Subtitle generation uses natural language processing techniques to extract important points and determine the display format based on them.
[0312] Step 4:
[0313] The terminal receives the generated subtitle data and displays it to the user in real time. The generated subtitle data is used as input, and the real-time display is output on the user's device. Libraries such as moviepy are used to perform the specific action of overlaying the subtitles onto the display screen.
[0314] Step 5:
[0315] Users visually review the subtitles presented to understand the content they are watching. They directly input the generated display information and receive real-time improvements to their viewing experience as output. This process involves using the displayed information to understand and process the viewing experience.
[0316] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0317] The system according to the present invention not only generates effective subtitles for videos, but also has the function of recognizing the user's emotions and reflecting them in the subtitles. This system mainly consists of a server, a terminal, and a user interface, and is realized through the following process.
[0318] The user uploads video data to the system using their own device. The server separates the audio from the uploaded video and converts it into text data using speech recognition technology. At this point, the text data is processed to generate concise subtitles.
[0319] Furthermore, this invention uses an emotion engine to recognize the user's emotions. This emotion recognition is performed by analyzing the user's facial expressions and voice tone through devices such as a camera and microphone connected to the terminal. For example, if the user shows a surprised expression, the emotion engine recognizes this and makes changes such as emphasizing the subtitle display style.
[0320] Based on the emotions it recognizes, the server dynamically adjusts the subtitle styling (e.g., changing colors and fonts) and display time, and further improves the accuracy of key point extraction for specific scenes. The subtitles generated in this way are output as a project file that can be freely edited with video editing software such as Final Cut Pro and Premiere Pro, and are provided to the user.
[0321] Users can import these project files into their video editing software and adjust them to best suit the context of their video work. In this way, the system of the present invention makes it possible to make the video viewing experience more interactive and engaging by utilizing emotional information.
[0322] The following describes the processing flow.
[0323] Step 1:
[0324] The user accesses the system interface from their terminal, selects the video data for which they want to generate subtitles, and uploads it to the server. The user interface provides a guide to simplify the process.
[0325] Step 2:
[0326] The server extracts audio from the uploaded video data. This process separates the audio track from the video file and converts it into a format suitable for analysis.
[0327] Step 3:
[0328] The server uses speech recognition technology to convert the extracted speech into text data. The spoken content is sequentially converted into text information through speech analysis.
[0329] Step 4:
[0330] The server extracts key points from the generated text data and produces clear subtitles. Natural language processing technology identifies relevant information and organizes it into a concise and easy-to-read format.
[0331] Step 5:
[0332] The device's built-in camera and microphone detect the user's facial expressions and voice tone, which are then analyzed by an emotion engine. For example, if the user smiles, that emotion is recognized as "joy."
[0333] Step 6:
[0334] The server adjusts subtitles based on the recognized user's emotions. By softening the color of the text or changing the font style, it provides readability that matches the user's emotions.
[0335] Step 7:
[0336] The server converts the final subtitle data into an editable project file format and provides the user with a download link. Users can download it to their own devices and make further adjustments using video editing software.
[0337] (Example 2)
[0338] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0339] When providing subtitles for video content, it is difficult to generate dynamic subtitle styles that reflect the viewer's emotional state, resulting in a challenge in adequately improving the immersion and quality of the viewing experience. Furthermore, there are limited ways to effectively incorporate user feedback.
[0340] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0341] In this invention, the server includes means for extracting audio from input video information, means for converting audio into text information using a speech recognition method, means for extracting key points from the text information and generating subtitles that highlight the key points, means for recognizing the user's emotions and dynamically adjusting the subtitle style based on the recognized emotions, and means for outputting the subtitle data in an editable format. This makes it possible to provide subtitles that reflect the user's emotional state, improve the quality of the viewing experience, and incorporate user feedback into the subtitle generation process.
[0342] "Input video information" refers to digital data, including visual and audio data, that the user provides to the system.
[0343] "Means for extracting audio" refers to a technology or device for separating and acquiring only audio data from video information.
[0344] "Speech recognition techniques" are algorithms and technologies used to convert speech data into strings of characters that a computer can understand.
[0345] "Textual information" refers to text data converted from audio data using speech recognition techniques, and is the information that forms the basis for subtitle generation.
[0346] "Means for extracting key points and generating concise subtitles" refers to algorithms or devices that analyze textual information, identify important content, and create subtitles based on that content.
[0347] "A means of recognizing user emotions and dynamically adjusting subtitle styles based on those emotions" refers to a technology that analyzes user emotions and changes the subtitle display format in real time based on the results.
[0348] "Means of outputting subtitle data in an editable format" refers to technologies and systems that provide generated subtitles in a file format that can be further edited using video editing software, etc.
[0349] This invention provides a system for generating emotion-based interactive subtitles for video content. This system primarily consists of a server, terminals, and a user interface.
[0350] Users can upload video data they want to view or edit to the system using their own devices. The server extracts audio from the received video data using multimedia tools and converts it into text using speech recognition technology. Common speech recognition APIs are used in this process.
[0351] Next, the server extracts key points from the text information and generates subtitles based on that content. A natural language processing model is used in this process. The generated subtitles are then styled in real time on the server according to the user's emotions, which are analyzed by an emotion engine via the camera and microphone on the user's device. For example, if the user smiles or shows a surprised expression, the color and font of the subtitles will change accordingly.
[0352] Furthermore, the server can output the generated subtitles as a project file that can be edited using the user's editing software. This allows users to freely add to and edit the provided subtitles to use them in a way that is more appropriate to the context of the video.
[0353] As a concrete example, imagine a scenario where a user uploads a travel video, and a system that detects the user's smile generates a caption saying "The joy of sharing this moment!" and changes the layout to a brighter color scheme.
[0354] An example of a prompt is, "Generate emotion-recognition-based subtitles for a video of our summer vacation together as a family." Using this prompt, users can naturally create subtitles that reflect their own emotions in the video.
[0355] Thus, the present invention provides subtitles that reflect the user's emotions, making the video viewing experience more interactive and engaging.
[0356] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0357] Step 1:
[0358] The user uploads video data to the system using their own device. The input for this step is video data, and the output is the transmission of data to the server. The user selects a file on the operation screen and sends that file to the system.
[0359] Step 2:
[0360] The server extracts audio from the uploaded video data. The input for this step is the video data sent by the user, and the output is the extracted audio data. The server uses multimedia processing tools to analyze the video data and perform the specific operation of separating the audio track.
[0361] Step 3:
[0362] The server converts the extracted audio data into text information using speech recognition technology. The input for this step is audio data, and the output is the converted text information. The server calls a speech recognition API to perform specific calculations that sequentially process the audio data and convert it into text.
[0363] Step 4:
[0364] The server extracts key points from text information to generate subtitles. The input for this step is text information, and the output is summarized subtitle data. The server utilizes natural language processing techniques to analyze the text, identify important key phrases, and perform the specific processing required to form subtitles.
[0365] Step 5:
[0366] The device recognizes the user's emotions and sends that data to the server. The input for this step is the user's facial expressions and voice tone collected from the device's camera and microphone, and the output is the recognized emotion data. The device uses emotion analysis software to perform specific calculations to determine the user's emotional state in real time.
[0367] Step 6:
[0368] The server dynamically adjusts the subtitle style based on the received sentiment data. The input for this step is the sentiment data and the generated subtitle data, and the output is the final, styled subtitle data. The server refers to style rules and performs specific operations to change the font, color, and display time of the subtitles.
[0369] Step 7:
[0370] The server outputs subtitle data in an editable format and provides it to the user. The input for this step is styled subtitle data, and the output is a project file usable with video editing software. The server uses a file format conversion tool to save the data in the appropriate format, making it available for user download.
[0371] (Application Example 2)
[0372] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0373] In recent years, as video content viewing has diversified, there has been a growing demand for interactive subtitles that respond to viewers' emotional states. However, conventional video subtitling systems have lacked the ability to reflect individual viewers' emotions, making it difficult to provide a deeper viewing experience. Furthermore, there has been a lack of subtitle editing technology that can reflect viewers' emotions in real time. Solving these problems is a challenge.
[0374] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0375] In this invention, the server includes means for extracting audio from input video data, means for converting audio into text data, and means for sensing emotions using an input device to analyze emotional states. This enables a more interactive and personalized viewing experience by dynamically adjusting subtitle styling and display time based on the viewer's emotional information.
[0376] An "input device" is a device used to acquire data such as audio and video and convert it into a format that can be used within the system.
[0377] "Speech recognition technology" is a technology that analyzes speech data and converts it into corresponding text data.
[0378] "Text data" refers to character information extracted from speech using speech recognition technology.
[0379] "Styling" refers to adjusting display attributes such as the color, font, and size of subtitles.
[0380] "Display characteristics" refer to visual features such as the color scheme and fonts of subtitles and interfaces.
[0381] "Emotion" refers to the psychological state, such as feelings and moods, that arise in the user's mind.
[0382] "Data format" refers to a standardized structure or format used when storing or processing digital information.
[0383] "Interaction" refers to the two-way interaction and exchange of information that takes place between a system and a user.
[0384] "Real-time" refers to a time concept where information is processed instantly, enabling immediate responses.
[0385] To realize this invention, three main components are necessary: a server, a terminal, and a user. The server extracts audio from video data and converts that audio into text data using speech recognition technology. Examples of speech recognition technologies used in this process include the Google Cloud Speech-to-Text API. Key points are extracted from the generated text data, and subtitles are dynamically generated.
[0386] The device plays a role in sensing the user's emotional state in real time. For this purpose, the device uses a built-in camera and microphone to analyze the user's facial expressions and tone of voice. The software used for analyzing emotional states utilizes the Azure Emotion Recognition API. This allows for the collection of real-time emotional data from the user, and the dynamic adjustment of subtitle styling and display time.
[0387] Users operate their own devices and upload video data to the server. The processed subtitles are provided in a data format that can be edited with video editing software such as Adobe Premiere Pro. In this way, the video viewing experience can be made more interactive and personalized.
[0388] For example, if a user is moved during the climax scene of a movie, the system can change the subtitle color to a softer shade that matches their emotion and even suggest an interaction such as, "Would you like to share this feeling with a friend?"
[0389] An example of a prompt sentence to input into a generative AI model is, "Please suggest a suitable subtitle style for this scene in this movie."
[0390] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0391] Step 1:
[0392] The user uploads video data from their device to the server. The uploaded video data is passed to the server as input. The server receives this data and prepares to analyze the video file.
[0393] Step 2:
[0394] The server extracts audio from the video data. Using the extracted audio data as input, it generates an audio file. This process transforms the video's audio information into a separate file, preparing it for the subsequent speech recognition process.
[0395] Step 3:
[0396] The server uses the Google Cloud Speech-to-Text API to convert the extracted audio data into text data. This text data is output, providing information in written form. Speech recognition technology transforms spoken content into readable text data.
[0397] Step 4:
[0398] The server analyzes the generated text data and applies an algorithm to extract key points. It extracts key points from the input text data and uses them to generate concise subtitle data that highlights the essentials. This results in a concise subtitle containing only the most important information.
[0399] Step 5:
[0400] The device senses the user's emotional state in real time using its camera and microphone. This sensor data is then used to analyze the user's emotional state using the Azure Emotion Recognition API. The analysis results are output, and the emotional state is quantified.
[0401] Step 6:
[0402] The server dynamically adjusts subtitle styling and display time based on user emotion data. Emotion data is used as input, and the color, font, and timing of subtitle display are changed in real time. This creates subtitles that visually reflect emotions.
[0403] Step 7:
[0404] The server generates adjusted subtitle data and outputs it in a format editable with Adobe Premiere Pro and other software. The output project file format is then provided to the user, allowing them to edit the video according to their own needs.
[0405] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0406] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0407] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0408] [Third Embodiment]
[0409] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0410] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0411] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0412] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0413] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0414] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0415] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0416] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0417] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0418] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0419] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0420] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0421] The system according to the present invention is for automatically generating effective subtitles for videos and providing them to users. This system consists of a server, a terminal, and a user interface.
[0422] First, the user uploads the video data for which they want subtitles to be generated to the system using their own device. The server extracts the audio from the received video data and uses speech recognition technology to convert that audio data into text data.
[0423] For example, in the case of a video recording of a lecture, the speaker's words are transcribed sequentially. The server then analyzes the generated text data and extracts the key points using natural language processing. This generates subtitles that highlight the important points of the lecture.
[0424] Furthermore, the server has a function to automatically identify speakers from audio data, and it is configured to display subtitles in different fonts and colors for each speaker. For example, in a video in a dialogue format, different speakers can be visually distinguished by applying blue and green fonts to each.
[0425] The generated subtitles are output in a project file format that can be edited with video editing software such as Final Cut Pro and Premiere Pro, according to the user's needs, allowing users to easily perform video editing tasks. Users can save the completed project file to their device and make any necessary adjustments to their final video work.
[0426] This system generates highly legible subtitles in a short time through an automated process, allowing users to achieve a professional finish without requiring specialized skills. Thus, the present invention aims to improve the efficiency of subtitle generation and editing in video production.
[0427] The following describes the processing flow.
[0428] Step 1:
[0429] The user uploads video data requiring subtitles to the server using their device. During this process, the user specifies the video file by dragging and dropping it using the system interface.
[0430] Step 2:
[0431] The server extracts the audio track from the received video data. This process separates the audio portion of the video file and saves it as an audio file.
[0432] Step 3:
[0433] The server analyzes the extracted audio using speech recognition technology and converts the audio data into text data. This conversion involves transcribing the content of the audio utterance by utterance, creating text from what was said.
[0434] Step 4:
[0435] The server analyzes the generated text data using natural language processing techniques and extracts the key points of the text to produce concise and clear subtitles. This creates subtitles that highlight particularly important information in the audio.
[0436] Step 5:
[0437] The server identifies the speaker from the audio data and sets different fonts and colors for each speaker. This identification makes it possible to visually distinguish between statements made by different speakers.
[0438] Step 6:
[0439] The server outputs the generated subtitle data as a project file format editable in Final Cut Pro or Premiere Pro. This project file contains the subtitle information and its styling.
[0440] Step 7:
[0441] Users download project files from the server and make final adjustments to subtitles using video editing software on their own devices. Within the editing environment, they can freely change the position and style of the subtitles and incorporate them into the final video.
[0442] (Example 1)
[0443] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0444] In current video production, generating visually clear subtitles requires considerable effort and specialized skills. Furthermore, differentiating subtitle display for each speaker and creating subtitles that capture the key points also takes time and effort, necessitating an efficient process.
[0445] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0446] In this invention, the server includes means for extracting audio information from video information, means for converting the audio information into text information using a speech recognition method, and means for extracting important information from the text information and generating display information that suppresses the important information. This makes it possible to quickly and effectively generate highly legible subtitles without requiring specialized knowledge.
[0447] "Video information" refers to information that includes video data and related audio data.
[0448] "Audio information" refers to audio data extracted from video information.
[0449] "Speech recognition method" is a technology that analyzes speech information and converts it into text information.
[0450] "Textual information" refers to text data obtained through speech recognition methods.
[0451] "Important information" refers to information that contains the main content or key points extracted from textual information.
[0452] "Display information" refers to subtitles and text display data generated based on important information.
[0453] "Speakers" refer to the individual individuals who are speaking within the audio information.
[0454] "Display format" refers to styles and formats used to visually change the appearance of text and graphics.
[0455] "Project information format" refers to a file format that saves display information in a format that allows for video editing.
[0456] "Rules" refer to established rules and standards for setting display formats based on speaker identification.
[0457] "Opinions" refers to feedback and evaluations from users.
[0458] This invention relates to a system for automatically generating highly legible subtitles from video information, and its main components are a user, a terminal, and a server. First, the user uploads video information requiring subtitles to the server using their own terminal. This includes file selection and upload operations via a web browser or a specific application.
[0459] The server analyzes the received video information and extracts the audio information. Audio processing libraries such as FFmpeg are used for this audio extraction. Next, the server applies a speech recognition method to convert the audio information into text information. Here, speech recognition technologies such as the Google Speech-to-Text API are used.
[0460] The converted text information is processed by a server, and important information is extracted by a natural language processing engine (e.g., spaCy). This generates display information containing the key points. The generated display information is then adjusted by speaker identification technology to output in a different display format for each speaker.
[0461] As a final result, the displayed information is output in project information format and can be downloaded to the user's device. Since this format is compatible with video editing software such as Final Cut Pro and Premiere Pro, users can further edit subtitles using this format.
[0462] As a concrete example, suppose a user uploads an educational video in which an instructor explains different topics. In this case, displaying the key points of each topic immediately after the start of each topic will make it easier for viewers to understand the content.
[0463] An example of a prompt to input into the generation AI model would be, "Generate the most relevant information for this educational video and visually highlight the key points of each section." This system allows users to quickly generate visually effective subtitles without requiring specialized knowledge or skills.
[0464] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0465] Step 1:
[0466] Users upload video information to the server using their own devices to generate subtitles. Specifically, users select video files and upload them using a dedicated application or web browser interface. The input is the video file, and the output is that file stored on the server.
[0467] Step 2:
[0468] The server extracts audio information from the uploaded video data. This extraction uses an audio processing library such as FFmpeg. This library extracts the audio track from the video file and writes it as an audio file. The input is a video file, and the output is a separated audio file.
[0469] Step 3:
[0470] The server uses speech recognition to convert extracted audio information into text. This process utilizes the Google Speech-to-Text API, sending audio data and receiving corresponding text data. The input is an audio file, and the output is text data.
[0471] Step 4:
[0472] The server applies natural language processing to the generated text information to extract important information. For example, it uses a natural language processing library such as spaCy to perform grammatical analysis and key phrase extraction. The input is text data, and the output is important information with key points highlighted.
[0473] Step 5:
[0474] The server uses speaker identification technology to adjust the display format of important information for each speaker. For example, it performs speech feature analysis and applies different fonts and colors based on each speaker's speech pattern. The input is important information, and the output is data displayed in a format appropriate for the speaker.
[0475] Step 6:
[0476] The server outputs the display information in project information format, making it available for download on the user's device. The project file is generated in a format that can be read by video editing software. The input is display data adjusted for each speaker, and the output is a project information file.
[0477] Users can download the project information file and perform further editing on their own devices.
[0478] (Application Example 1)
[0479] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0480] In modern content distribution services, viewers watch videos in a variety of situations, and there is a growing need for automatic subtitle generation to aid visual comprehension, especially for videos with difficult-to-understand audio or in different languages. Therefore, there is a demand for real-time, high-quality, and visually clear subtitle generation to improve the viewing experience.
[0481] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0482] In this invention, the server includes means for extracting audio from input video data, means for converting audio into text data using speech recognition technology, and control means for generating and visually presenting display information in real time. This allows viewers to visually obtain high-quality subtitles that are automatically generated in real time, making it easier to understand the content they are watching.
[0483] "Video data" refers to digital data that includes visual information, including dynamic screen information such as videos.
[0484] "Speech recognition technology" is a technology that analyzes collected speech data and converts it into text information based on that analysis.
[0485] "Text data" refers to text information converted by speech recognition technology, and is data expressed in a form that is readable by humans.
[0486] "Display information" refers to data presented visually, particularly subtitles and supplementary information displayed within a video.
[0487] "Speaker identification" is the process of identifying and distinguishing different speakers, and it is performed using the characteristics of voices and sounds within the video.
[0488] "Project data" refers to a format of editable data used in video editing, and includes a set of data files containing subtitles and other video editing elements.
[0489] "Control means" refers to mechanisms and programs for managing and operating various processes in an image processing system, enabling real-time processing.
[0490] "Real-time" refers to a temporal concept where instructed actions or processes are carried out immediately without delay, signifying an immediate response.
[0491] This invention provides a real-time subtitle generation system for a video distribution platform. This system consists of a server, a terminal, and a user interface.
[0492] The server first receives video data for the user to view. When extracting audio from the video data, it uses speech recognition technology to convert the audio into text data. For example, the Python `speech_recognition` library can be used to perform this audio conversion efficiently.
[0493] The converted text data is analyzed on the server and visually presented as real-time display information. This includes compositing subtitles onto the video using video editing software such as the moviepy library. Each speaker's statements are displayed in a different color and font, making it easy for viewers to distinguish between speakers.
[0494] The device instantly visualizes the generated display information on the user's device, allowing the user to view high-quality subtitles in real time while watching. This makes it easier to understand the content being viewed.
[0495] For example, when a user is watching a cooking video, automatically generated subtitles are displayed by the system, allowing viewers to more clearly understand the steps of the recipe. To further enhance user convenience, the subtitle format and display style can be customized through the device settings.
[0496] An example of a prompt message is: "Generate subtitles for a cooking video. Create subtitles that clearly indicate ingredients and steps, and visually distinguish between different speakers (chef and narrator)." This prompt allows users to easily generate display information tailored to their specific needs.
[0497] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0498] Step 1:
[0499] The server receives video data being streamed from the user and extracts audio from that data. This process uses video data as input and outputs an audio stream. The extraction of audio data involves using the necessary APIs and libraries to obtain the audio file.
[0500] Step 2:
[0501] The server uses speech recognition technology to convert the extracted audio stream into text data. This process takes the audio stream as input and obtains text data as output. The speech_recognition library is used to perform speech analysis and convert the spoken portion into text.
[0502] Step 3:
[0503] The server analyzes the converted text data and generates subtitles as display information. Here, text data is used as input, and subtitles in a visually displayable format are output. Subtitle generation uses natural language processing techniques to extract important points and determine the display format based on them.
[0504] Step 4:
[0505] The terminal receives the generated subtitle data and displays it to the user in real time. The generated subtitle data is used as input, and the real-time display is output on the user's device. Libraries such as moviepy are used to perform the specific action of overlaying the subtitles onto the display screen.
[0506] Step 5:
[0507] Users visually review the subtitles presented to understand the content they are watching. They directly input the generated display information and receive real-time improvements to their viewing experience as output. This process involves using the displayed information to understand and process the viewing experience.
[0508] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0509] The system according to the present invention not only generates effective subtitles for videos, but also has the function of recognizing the user's emotions and reflecting them in the subtitles. This system mainly consists of a server, a terminal, and a user interface, and is realized through the following process.
[0510] The user uploads video data to the system using their own device. The server separates the audio from the uploaded video and converts it into text data using speech recognition technology. At this point, the text data is processed to generate concise subtitles.
[0511] Furthermore, this invention uses an emotion engine to recognize the user's emotions. This emotion recognition is performed by analyzing the user's facial expressions and voice tone through devices such as a camera and microphone connected to the terminal. For example, if the user shows a surprised expression, the emotion engine recognizes this and makes changes such as emphasizing the subtitle display style.
[0512] Based on the emotions it recognizes, the server dynamically adjusts the subtitle styling (e.g., changing colors and fonts) and display time, and further improves the accuracy of key point extraction for specific scenes. The subtitles generated in this way are output as a project file that can be freely edited with video editing software such as Final Cut Pro and Premiere Pro, and are provided to the user.
[0513] Users can import these project files into their video editing software and adjust them to best suit the context of their video work. In this way, the system of the present invention makes it possible to make the video viewing experience more interactive and engaging by utilizing emotional information.
[0514] The following describes the processing flow.
[0515] Step 1:
[0516] The user accesses the system interface from their terminal, selects the video data for which they want to generate subtitles, and uploads it to the server. The user interface provides a guide to simplify the process.
[0517] Step 2:
[0518] The server extracts audio from the uploaded video data. This process separates the audio track from the video file and converts it into a format suitable for analysis.
[0519] Step 3:
[0520] The server uses speech recognition technology to convert the extracted speech into text data. The spoken content is sequentially converted into text information through speech analysis.
[0521] Step 4:
[0522] The server extracts key points from the generated text data and produces clear subtitles. Natural language processing technology identifies relevant information and organizes it into a concise and easy-to-read format.
[0523] Step 5:
[0524] The device's built-in camera and microphone detect the user's facial expressions and voice tone, which are then analyzed by an emotion engine. For example, if the user smiles, that emotion is recognized as "joy."
[0525] Step 6:
[0526] The server adjusts subtitles based on the recognized user's emotions. By softening the color of the text or changing the font style, it provides readability that matches the user's emotions.
[0527] Step 7:
[0528] The server converts the final subtitle data into an editable project file format and provides the user with a download link. Users can download it to their own devices and make further adjustments using video editing software.
[0529] (Example 2)
[0530] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0531] When providing subtitles for video content, it is difficult to generate dynamic subtitle styles that reflect the viewer's emotional state, resulting in a challenge in adequately improving the immersion and quality of the viewing experience. Furthermore, there are limited ways to effectively incorporate user feedback.
[0532] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0533] In this invention, the server includes means for extracting audio from input video information, means for converting audio into text information using a speech recognition method, means for extracting key points from the text information and generating subtitles that highlight the key points, means for recognizing the user's emotions and dynamically adjusting the subtitle style based on the recognized emotions, and means for outputting the subtitle data in an editable format. This makes it possible to provide subtitles that reflect the user's emotional state, improve the quality of the viewing experience, and incorporate user feedback into the subtitle generation process.
[0534] "Input video information" refers to digital data, including visual and audio data, that the user provides to the system.
[0535] "Means for extracting audio" refers to a technology or device for separating and acquiring only audio data from video information.
[0536] "Speech recognition techniques" are algorithms and technologies used to convert speech data into strings of characters that a computer can understand.
[0537] "Textual information" refers to text data converted from audio data using speech recognition techniques, and is the information that forms the basis for subtitle generation.
[0538] "Means for extracting key points and generating concise subtitles" refers to algorithms or devices that analyze textual information, identify important content, and create subtitles based on that content.
[0539] "A means of recognizing user emotions and dynamically adjusting subtitle styles based on those emotions" refers to a technology that analyzes user emotions and changes the subtitle display format in real time based on the results.
[0540] "Means of outputting subtitle data in an editable format" refers to technologies and systems that provide generated subtitles in a file format that can be further edited using video editing software, etc.
[0541] This invention provides a system for generating emotion-based interactive subtitles for video content. This system primarily consists of a server, terminals, and a user interface.
[0542] Users can upload video data they want to view or edit to the system using their own devices. The server extracts audio from the received video data using multimedia tools and converts it into text using speech recognition technology. Common speech recognition APIs are used in this process.
[0543] Next, the server extracts key points from the text information and generates subtitles based on that content. A natural language processing model is used in this process. The generated subtitles are then styled in real time on the server according to the user's emotions, which are analyzed by an emotion engine via the camera and microphone on the user's device. For example, if the user smiles or shows a surprised expression, the color and font of the subtitles will change accordingly.
[0544] Furthermore, the server can output the generated subtitles as a project file that can be edited using the user's editing software. This allows users to freely add to and edit the provided subtitles to use them in a way that is more appropriate to the context of the video.
[0545] As a concrete example, imagine a scenario where a user uploads a travel video, and a system that detects the user's smile generates a caption saying "The joy of sharing this moment!" and changes the layout to a brighter color scheme.
[0546] An example of a prompt is, "Generate emotion-recognition-based subtitles for a video of our summer vacation together as a family." Using this prompt, users can naturally create subtitles that reflect their own emotions in the video.
[0547] Thus, the present invention provides subtitles that reflect the user's emotions, making the video viewing experience more interactive and engaging.
[0548] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0549] Step 1:
[0550] The user uploads video data to the system using their own device. The input for this step is video data, and the output is the transmission of data to the server. The user selects a file on the operation screen and sends that file to the system.
[0551] Step 2:
[0552] The server extracts audio from the uploaded video data. The input for this step is the video data sent by the user, and the output is the extracted audio data. The server uses multimedia processing tools to analyze the video data and perform the specific operation of separating the audio track.
[0553] Step 3:
[0554] The server converts the extracted audio data into text information using speech recognition technology. The input for this step is audio data, and the output is the converted text information. The server calls a speech recognition API to perform specific calculations that sequentially process the audio data and convert it into text.
[0555] Step 4:
[0556] The server extracts key points from text information to generate subtitles. The input for this step is text information, and the output is summarized subtitle data. The server utilizes natural language processing techniques to analyze the text, identify important key phrases, and perform the specific processing required to form subtitles.
[0557] Step 5:
[0558] The device recognizes the user's emotions and sends that data to the server. The input for this step is the user's facial expressions and voice tone collected from the device's camera and microphone, and the output is the recognized emotion data. The device uses emotion analysis software to perform specific calculations to determine the user's emotional state in real time.
[0559] Step 6:
[0560] The server dynamically adjusts the subtitle style based on the received sentiment data. The input for this step is the sentiment data and the generated subtitle data, and the output is the final, styled subtitle data. The server refers to style rules and performs specific operations to change the font, color, and display time of the subtitles.
[0561] Step 7:
[0562] The server outputs subtitle data in an editable format and provides it to the user. The input for this step is styled subtitle data, and the output is a project file usable with video editing software. The server uses a file format conversion tool to save the data in the appropriate format, making it available for user download.
[0563] (Application Example 2)
[0564] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0565] In recent years, as video content viewing has diversified, there has been a growing demand for interactive subtitles that respond to viewers' emotional states. However, conventional video subtitling systems have lacked the ability to reflect individual viewers' emotions, making it difficult to provide a deeper viewing experience. Furthermore, there has been a lack of subtitle editing technology that can reflect viewers' emotions in real time. Solving these problems is a challenge.
[0566] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0567] In this invention, the server includes means for extracting audio from input video data, means for converting audio into text data, and means for sensing emotions using an input device to analyze emotional states. This enables a more interactive and personalized viewing experience by dynamically adjusting subtitle styling and display time based on the viewer's emotional information.
[0568] An "input device" is a device used to acquire data such as audio and video and convert it into a format that can be used within the system.
[0569] "Speech recognition technology" is a technology that analyzes speech data and converts it into corresponding text data.
[0570] "Text data" refers to character information extracted from speech using speech recognition technology.
[0571] "Styling" refers to adjusting display attributes such as the color, font, and size of subtitles.
[0572] "Display characteristics" refer to visual features such as the color scheme and fonts of subtitles and interfaces.
[0573] "Emotion" refers to the psychological state, such as feelings and moods, that arise in the user's mind.
[0574] "Data format" refers to a standardized structure or format used when storing or processing digital information.
[0575] "Interaction" refers to the two-way interaction and exchange of information that takes place between a system and a user.
[0576] "Real-time" refers to a time concept where information is processed instantly, enabling immediate responses.
[0577] To realize this invention, three main components are necessary: a server, a terminal, and a user. The server extracts audio from video data and converts that audio into text data using speech recognition technology. Examples of speech recognition technologies used in this process include the Google Cloud Speech-to-Text API. Key points are extracted from the generated text data, and subtitles are dynamically generated.
[0578] The device plays a role in sensing the user's emotional state in real time. For this purpose, the device uses a built-in camera and microphone to analyze the user's facial expressions and tone of voice. The software used for analyzing emotional states utilizes the Azure Emotion Recognition API. This allows for the collection of real-time emotional data from the user, and the dynamic adjustment of subtitle styling and display time.
[0579] Users operate their own devices and upload video data to the server. The processed subtitles are provided in a data format that can be edited with video editing software such as Adobe Premiere Pro. In this way, the video viewing experience can be made more interactive and personalized.
[0580] For example, if a user is moved during the climax scene of a movie, the system can change the subtitle color to a softer shade that matches their emotion and even suggest an interaction such as, "Would you like to share this feeling with a friend?"
[0581] An example of a prompt sentence to input into a generative AI model is, "Please suggest a suitable subtitle style for this scene in this movie."
[0582] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0583] Step 1:
[0584] The user uploads video data from their device to the server. The uploaded video data is passed to the server as input. The server receives this data and prepares to analyze the video file.
[0585] Step 2:
[0586] The server extracts audio from the video data. Using the extracted audio data as input, it generates an audio file. This process transforms the video's audio information into a separate file, preparing it for the subsequent speech recognition process.
[0587] Step 3:
[0588] The server uses the Google Cloud Speech-to-Text API to convert the extracted audio data into text data. This text data is output, providing information in written form. Speech recognition technology transforms spoken content into readable text data.
[0589] Step 4:
[0590] The server analyzes the generated text data and applies an algorithm to extract key points. It extracts key points from the input text data and uses them to generate concise subtitle data that highlights the essentials. This results in a concise subtitle containing only the most important information.
[0591] Step 5:
[0592] The device senses the user's emotional state in real time using its camera and microphone. This sensor data is then used to analyze the user's emotional state using the Azure Emotion Recognition API. The analysis results are output, and the emotional state is quantified.
[0593] Step 6:
[0594] The server dynamically adjusts subtitle styling and display time based on user emotion data. Emotion data is used as input, and the color, font, and timing of subtitle display are changed in real time. This creates subtitles that visually reflect emotions.
[0595] Step 7:
[0596] The server generates adjusted subtitle data and outputs it in a format editable with Adobe Premiere Pro and other software. The output project file format is then provided to the user, allowing them to edit the video according to their own needs.
[0597] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0598] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0599] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0600] [Fourth Embodiment]
[0601] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0602] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0603] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0604] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0605] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0606] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0607] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0608] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0609] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0610] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0611] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0612] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0613] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0614] The system according to the present invention is for automatically generating effective subtitles for videos and providing them to users. This system consists of a server, a terminal, and a user interface.
[0615] First, the user uploads the video data for which they want subtitles to be generated to the system using their own device. The server extracts the audio from the received video data and uses speech recognition technology to convert that audio data into text data.
[0616] For example, in the case of a video recording of a lecture, the speaker's words are transcribed sequentially. The server then analyzes the generated text data and extracts the key points using natural language processing. This generates subtitles that highlight the important points of the lecture.
[0617] Furthermore, the server has a function to automatically identify speakers from audio data, and it is configured to display subtitles in different fonts and colors for each speaker. For example, in a video in a dialogue format, different speakers can be visually distinguished by applying blue and green fonts to each.
[0618] The generated subtitles are output in a project file format that can be edited with video editing software such as Final Cut Pro and Premiere Pro, according to the user's needs, allowing users to easily perform video editing tasks. Users can save the completed project file to their device and make any necessary adjustments to their final video work.
[0619] This system generates highly legible subtitles in a short time through an automated process, allowing users to achieve a professional finish without requiring specialized skills. Thus, the present invention aims to improve the efficiency of subtitle generation and editing in video production.
[0620] The following describes the processing flow.
[0621] Step 1:
[0622] The user uploads video data requiring subtitles to the server using their device. During this process, the user specifies the video file by dragging and dropping it using the system interface.
[0623] Step 2:
[0624] The server extracts the audio track from the received video data. This process separates the audio portion of the video file and saves it as an audio file.
[0625] Step 3:
[0626] The server analyzes the extracted audio using speech recognition technology and converts the audio data into text data. This conversion involves transcribing the content of the audio utterance by utterance, creating text from what was said.
[0627] Step 4:
[0628] The server analyzes the generated text data using natural language processing techniques and extracts the key points of the text to produce concise and clear subtitles. This creates subtitles that highlight particularly important information in the audio.
[0629] Step 5:
[0630] The server identifies the speaker from the audio data and sets different fonts and colors for each speaker. This identification makes it possible to visually distinguish between statements made by different speakers.
[0631] Step 6:
[0632] The server outputs the generated subtitle data as a project file format editable in Final Cut Pro or Premiere Pro. This project file contains the subtitle information and its styling.
[0633] Step 7:
[0634] Users download project files from the server and make final adjustments to subtitles using video editing software on their own devices. Within the editing environment, they can freely change the position and style of the subtitles and incorporate them into the final video.
[0635] (Example 1)
[0636] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0637] In current video production, generating visually clear subtitles requires considerable effort and specialized skills. Furthermore, differentiating subtitle display for each speaker and creating subtitles that capture the key points also takes time and effort, necessitating an efficient process.
[0638] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0639] In this invention, the server includes means for extracting audio information from video information, means for converting the audio information into text information using a speech recognition method, and means for extracting important information from the text information and generating display information that suppresses the important information. This makes it possible to quickly and effectively generate highly legible subtitles without requiring specialized knowledge.
[0640] "Video information" refers to information that includes video data and related audio data.
[0641] "Audio information" refers to audio data extracted from video information.
[0642] "Speech recognition method" is a technology that analyzes speech information and converts it into text information.
[0643] "Textual information" refers to text data obtained through speech recognition methods.
[0644] "Important information" refers to information that contains the main content or key points extracted from textual information.
[0645] "Display information" refers to subtitles and text display data generated based on important information.
[0646] "Speakers" refer to the individual individuals who are speaking within the audio information.
[0647] "Display format" refers to styles and formats used to visually change the appearance of text and graphics.
[0648] "Project information format" refers to a file format that saves display information in a format that allows for video editing.
[0649] "Rules" refer to established rules and standards for setting display formats based on speaker identification.
[0650] "Opinions" refers to feedback and evaluations from users.
[0651] This invention relates to a system for automatically generating highly legible subtitles from video information, and its main components are a user, a terminal, and a server. First, the user uploads video information requiring subtitles to the server using their own terminal. This includes file selection and upload operations via a web browser or a specific application.
[0652] The server analyzes the received video information and extracts the audio information. Audio processing libraries such as FFmpeg are used for this audio extraction. Next, the server applies a speech recognition method to convert the audio information into text information. Here, speech recognition technologies such as the Google Speech-to-Text API are used.
[0653] The converted text information is processed by a server, and important information is extracted by a natural language processing engine (e.g., spaCy). This generates display information containing the key points. The generated display information is then adjusted by speaker identification technology to output in a different display format for each speaker.
[0654] As a final result, the displayed information is output in project information format and can be downloaded to the user's device. Since this format is compatible with video editing software such as Final Cut Pro and Premiere Pro, users can further edit subtitles using this format.
[0655] As a concrete example, suppose a user uploads an educational video in which an instructor explains different topics. In this case, displaying the key points of each topic immediately after the start of each topic will make it easier for viewers to understand the content.
[0656] An example of a prompt to input into the generation AI model would be, "Generate the most relevant information for this educational video and visually highlight the key points of each section." This system allows users to quickly generate visually effective subtitles without requiring specialized knowledge or skills.
[0657] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0658] Step 1:
[0659] Users upload video information to the server using their own devices to generate subtitles. Specifically, users select video files and upload them using a dedicated application or web browser interface. The input is the video file, and the output is that file stored on the server.
[0660] Step 2:
[0661] The server extracts audio information from the uploaded video data. This extraction uses an audio processing library such as FFmpeg. This library extracts the audio track from the video file and writes it as an audio file. The input is a video file, and the output is a separated audio file.
[0662] Step 3:
[0663] The server uses speech recognition to convert extracted audio information into text. This process utilizes the Google Speech-to-Text API, sending audio data and receiving corresponding text data. The input is an audio file, and the output is text data.
[0664] Step 4:
[0665] The server applies natural language processing to the generated text information to extract important information. For example, it uses a natural language processing library such as spaCy to perform grammatical analysis and key phrase extraction. The input is text data, and the output is important information with key points highlighted.
[0666] Step 5:
[0667] The server uses speaker identification technology to adjust the display format of important information for each speaker. For example, it performs speech feature analysis and applies different fonts and colors based on each speaker's speech pattern. The input is important information, and the output is data displayed in a format appropriate for the speaker.
[0668] Step 6:
[0669] The server outputs the display information in project information format, making it available for download on the user's device. The project file is generated in a format that can be read by video editing software. The input is display data adjusted for each speaker, and the output is a project information file.
[0670] Users can download the project information file and perform further editing on their own devices.
[0671] (Application Example 1)
[0672] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0673] In modern content distribution services, viewers watch videos in a variety of situations, and there is a growing need for automatic subtitle generation to aid visual comprehension, especially for videos with difficult-to-understand audio or in different languages. Therefore, there is a demand for real-time, high-quality, and visually clear subtitle generation to improve the viewing experience.
[0674] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0675] In this invention, the server includes means for extracting audio from input video data, means for converting audio into text data using speech recognition technology, and control means for generating and visually presenting display information in real time. This allows viewers to visually obtain high-quality subtitles that are automatically generated in real time, making it easier to understand the content they are watching.
[0676] "Video data" refers to digital data that includes visual information, including dynamic screen information such as videos.
[0677] "Speech recognition technology" is a technology that analyzes collected speech data and converts it into text information based on that analysis.
[0678] "Text data" refers to text information converted by speech recognition technology, and is data expressed in a form that is readable by humans.
[0679] "Display information" refers to data presented visually, particularly subtitles and supplementary information displayed within a video.
[0680] "Speaker identification" is the process of identifying and distinguishing different speakers, and it is performed using the characteristics of voices and sounds within the video.
[0681] "Project data" refers to a format of editable data used in video editing, and includes a set of data files containing subtitles and other video editing elements.
[0682] "Control means" refers to mechanisms and programs for managing and operating various processes in an image processing system, enabling real-time processing.
[0683] "Real-time" refers to a temporal concept where instructed actions or processes are carried out immediately without delay, signifying an immediate response.
[0684] This invention provides a real-time subtitle generation system for a video distribution platform. This system consists of a server, a terminal, and a user interface.
[0685] The server first receives video data for the user to view. When extracting audio from the video data, it uses speech recognition technology to convert the audio into text data. For example, the Python `speech_recognition` library can be used to perform this audio conversion efficiently.
[0686] The converted text data is analyzed on the server and visually presented as real-time display information. This includes compositing subtitles onto the video using video editing software such as the moviepy library. Each speaker's statements are displayed in a different color and font, making it easy for viewers to distinguish between speakers.
[0687] The device instantly visualizes the generated display information on the user's device, allowing the user to view high-quality subtitles in real time while watching. This makes it easier to understand the content being viewed.
[0688] For example, when a user is watching a cooking video, automatically generated subtitles are displayed by the system, allowing viewers to more clearly understand the steps of the recipe. To further enhance user convenience, the subtitle format and display style can be customized through the device settings.
[0689] An example of a prompt message is: "Generate subtitles for a cooking video. Create subtitles that clearly indicate ingredients and steps, and visually distinguish between different speakers (chef and narrator)." This prompt allows users to easily generate display information tailored to their specific needs.
[0690] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0691] Step 1:
[0692] The server receives video data being streamed from the user and extracts audio from that data. This process uses video data as input and outputs an audio stream. The extraction of audio data involves using the necessary APIs and libraries to obtain the audio file.
[0693] Step 2:
[0694] The server uses speech recognition technology to convert the extracted audio stream into text data. This process takes the audio stream as input and obtains text data as output. The speech_recognition library is used to perform speech analysis and convert the spoken portion into text.
[0695] Step 3:
[0696] The server analyzes the converted text data and generates subtitles as display information. Here, text data is used as input, and subtitles in a visually displayable format are output. Subtitle generation uses natural language processing techniques to extract important points and determine the display format based on them.
[0697] Step 4:
[0698] The terminal receives the generated subtitle data and displays it to the user in real time. The generated subtitle data is used as input, and the real-time display is output on the user's device. Libraries such as moviepy are used to perform the specific action of overlaying the subtitles onto the display screen.
[0699] Step 5:
[0700] Users visually review the subtitles presented to understand the content they are watching. They directly input the generated display information and receive real-time improvements to their viewing experience as output. This process involves using the displayed information to understand and process the viewing experience.
[0701] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0702] The system according to the present invention not only generates effective subtitles for videos, but also has the function of recognizing the user's emotions and reflecting them in the subtitles. This system mainly consists of a server, a terminal, and a user interface, and is realized through the following process.
[0703] The user uploads video data to the system using their own device. The server separates the audio from the uploaded video and converts it into text data using speech recognition technology. At this point, the text data is processed to generate concise subtitles.
[0704] Furthermore, this invention uses an emotion engine to recognize the user's emotions. This emotion recognition is performed by analyzing the user's facial expressions and voice tone through devices such as a camera and microphone connected to the terminal. For example, if the user shows a surprised expression, the emotion engine recognizes this and makes changes such as emphasizing the subtitle display style.
[0705] Based on the emotions it recognizes, the server dynamically adjusts the subtitle styling (e.g., changing colors and fonts) and display time, and further improves the accuracy of key point extraction for specific scenes. The subtitles generated in this way are output as a project file that can be freely edited with video editing software such as Final Cut Pro and Premiere Pro, and are provided to the user.
[0706] Users can import these project files into their video editing software and adjust them to best suit the context of their video work. In this way, the system of the present invention makes it possible to make the video viewing experience more interactive and engaging by utilizing emotional information.
[0707] The following describes the processing flow.
[0708] Step 1:
[0709] The user accesses the system interface from their terminal, selects the video data for which they want to generate subtitles, and uploads it to the server. The user interface provides a guide to simplify the process.
[0710] Step 2:
[0711] The server extracts audio from the uploaded video data. This process separates the audio track from the video file and converts it into a format suitable for analysis.
[0712] Step 3:
[0713] The server uses speech recognition technology to convert the extracted speech into text data. The spoken content is sequentially converted into text information through speech analysis.
[0714] Step 4:
[0715] The server extracts key points from the generated text data and produces clear subtitles. Natural language processing technology identifies relevant information and organizes it into a concise and easy-to-read format.
[0716] Step 5:
[0717] The device's built-in camera and microphone detect the user's facial expressions and voice tone, which are then analyzed by an emotion engine. For example, if the user smiles, that emotion is recognized as "joy."
[0718] Step 6:
[0719] The server adjusts subtitles based on the recognized user's emotions. By softening the color of the text or changing the font style, it provides readability that matches the user's emotions.
[0720] Step 7:
[0721] The server converts the final subtitle data into an editable project file format and provides the user with a download link. Users can download it to their own devices and make further adjustments using video editing software.
[0722] (Example 2)
[0723] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0724] When providing subtitles for video content, it is difficult to generate dynamic subtitle styles that reflect the viewer's emotional state, resulting in a challenge in adequately improving the immersion and quality of the viewing experience. Furthermore, there are limited ways to effectively incorporate user feedback.
[0725] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0726] In this invention, the server includes means for extracting audio from input video information, means for converting audio into text information using a speech recognition method, means for extracting key points from the text information and generating subtitles that highlight the key points, means for recognizing the user's emotions and dynamically adjusting the subtitle style based on the recognized emotions, and means for outputting the subtitle data in an editable format. This makes it possible to provide subtitles that reflect the user's emotional state, improve the quality of the viewing experience, and incorporate user feedback into the subtitle generation process.
[0727] "Input video information" refers to digital data, including visual and audio data, that the user provides to the system.
[0728] "Means for extracting audio" refers to a technology or device for separating and acquiring only audio data from video information.
[0729] "Speech recognition techniques" are algorithms and technologies used to convert speech data into strings of characters that a computer can understand.
[0730] "Textual information" refers to text data converted from audio data using speech recognition techniques, and is the information that forms the basis for subtitle generation.
[0731] "Means for extracting key points and generating concise subtitles" refers to algorithms or devices that analyze textual information, identify important content, and create subtitles based on that content.
[0732] "A means of recognizing user emotions and dynamically adjusting subtitle styles based on those emotions" refers to a technology that analyzes user emotions and changes the subtitle display format in real time based on the results.
[0733] "Means of outputting subtitle data in an editable format" refers to technologies and systems that provide generated subtitles in a file format that can be further edited using video editing software, etc.
[0734] This invention provides a system for generating emotion-based interactive subtitles for video content. This system primarily consists of a server, terminals, and a user interface.
[0735] Users can upload video data they want to view or edit to the system using their own devices. The server extracts audio from the received video data using multimedia tools and converts it into text using speech recognition technology. Common speech recognition APIs are used in this process.
[0736] Next, the server extracts key points from the text information and generates subtitles based on that content. A natural language processing model is used in this process. The generated subtitles are then styled in real time on the server according to the user's emotions, which are analyzed by an emotion engine via the camera and microphone on the user's device. For example, if the user smiles or shows a surprised expression, the color and font of the subtitles will change accordingly.
[0737] Furthermore, the server can output the generated subtitles as a project file that can be edited using the user's editing software. This allows users to freely add to and edit the provided subtitles to use them in a way that is more appropriate to the context of the video.
[0738] As a concrete example, imagine a scenario where a user uploads a travel video, and a system that detects the user's smile generates a caption saying "The joy of sharing this moment!" and changes the layout to a brighter color scheme.
[0739] An example of a prompt is, "Generate emotion-recognition-based subtitles for a video of our summer vacation together as a family." Using this prompt, users can naturally create subtitles that reflect their own emotions in the video.
[0740] Thus, the present invention provides subtitles that reflect the user's emotions, making the video viewing experience more interactive and engaging.
[0741] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0742] Step 1:
[0743] The user uploads video data to the system using their own device. The input for this step is video data, and the output is the transmission of data to the server. The user selects a file on the operation screen and sends that file to the system.
[0744] Step 2:
[0745] The server extracts audio from the uploaded video data. The input for this step is the video data sent by the user, and the output is the extracted audio data. The server uses multimedia processing tools to analyze the video data and perform the specific operation of separating the audio track.
[0746] Step 3:
[0747] The server converts the extracted audio data into text information using speech recognition technology. The input for this step is audio data, and the output is the converted text information. The server calls a speech recognition API to perform specific calculations that sequentially process the audio data and convert it into text.
[0748] Step 4:
[0749] The server extracts key points from text information to generate subtitles. The input for this step is text information, and the output is summarized subtitle data. The server utilizes natural language processing techniques to analyze the text, identify important key phrases, and perform the specific processing required to form subtitles.
[0750] Step 5:
[0751] The device recognizes the user's emotions and sends that data to the server. The input for this step is the user's facial expressions and voice tone collected from the device's camera and microphone, and the output is the recognized emotion data. The device uses emotion analysis software to perform specific calculations to determine the user's emotional state in real time.
[0752] Step 6:
[0753] The server dynamically adjusts the subtitle style based on the received sentiment data. The input for this step is the sentiment data and the generated subtitle data, and the output is the final, styled subtitle data. The server refers to style rules and performs specific operations to change the font, color, and display time of the subtitles.
[0754] Step 7:
[0755] The server outputs subtitle data in an editable format and provides it to the user. The input for this step is styled subtitle data, and the output is a project file usable with video editing software. The server uses a file format conversion tool to save the data in the appropriate format, making it available for user download.
[0756] (Application Example 2)
[0757] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0758] In recent years, as video content viewing has diversified, there has been a growing demand for interactive subtitles that respond to viewers' emotional states. However, conventional video subtitling systems have lacked the ability to reflect individual viewers' emotions, making it difficult to provide a deeper viewing experience. Furthermore, there has been a lack of subtitle editing technology that can reflect viewers' emotions in real time. Solving these problems is a challenge.
[0759] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0760] In this invention, the server includes means for extracting audio from input video data, means for converting audio into text data, and means for sensing emotions using an input device to analyze emotional states. This enables a more interactive and personalized viewing experience by dynamically adjusting subtitle styling and display time based on the viewer's emotional information.
[0761] An "input device" is a device used to acquire data such as audio and video and convert it into a format that can be used within the system.
[0762] "Speech recognition technology" is a technology that analyzes speech data and converts it into corresponding text data.
[0763] "Text data" refers to character information extracted from speech using speech recognition technology.
[0764] "Styling" refers to adjusting display attributes such as the color, font, and size of subtitles.
[0765] "Display characteristics" refer to visual features such as the color scheme and fonts of subtitles and interfaces.
[0766] "Emotion" refers to the psychological state, such as feelings and moods, that arise in the user's mind.
[0767] "Data format" refers to a standardized structure or format used when storing or processing digital information.
[0768] "Interaction" refers to the two-way interaction and exchange of information that takes place between a system and a user.
[0769] "Real-time" refers to a time concept where information is processed instantly, enabling immediate responses.
[0770] To realize this invention, three main components are necessary: a server, a terminal, and a user. The server extracts audio from video data and converts that audio into text data using speech recognition technology. Examples of speech recognition technologies used in this process include the Google Cloud Speech-to-Text API. Key points are extracted from the generated text data, and subtitles are dynamically generated.
[0771] The device plays a role in sensing the user's emotional state in real time. For this purpose, the device uses a built-in camera and microphone to analyze the user's facial expressions and tone of voice. The software used for analyzing emotional states utilizes the Azure Emotion Recognition API. This allows for the collection of real-time emotional data from the user, and the dynamic adjustment of subtitle styling and display time.
[0772] Users operate their own devices and upload video data to the server. The processed subtitles are provided in a data format that can be edited with video editing software such as Adobe Premiere Pro. In this way, the video viewing experience can be made more interactive and personalized.
[0773] For example, if a user is moved during the climax scene of a movie, the system can change the subtitle color to a softer shade that matches their emotion and even suggest an interaction such as, "Would you like to share this feeling with a friend?"
[0774] An example of a prompt sentence to input into a generative AI model is, "Please suggest a suitable subtitle style for this scene in this movie."
[0775] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0776] Step 1:
[0777] The user uploads video data from their device to the server. The uploaded video data is passed to the server as input. The server receives this data and prepares to analyze the video file.
[0778] Step 2:
[0779] The server extracts audio from the video data. Using the extracted audio data as input, it generates an audio file. This process transforms the video's audio information into a separate file, preparing it for the subsequent speech recognition process.
[0780] Step 3:
[0781] The server uses the Google Cloud Speech-to-Text API to convert the extracted audio data into text data. This text data is output, providing information in written form. Speech recognition technology transforms spoken content into readable text data.
[0782] Step 4:
[0783] The server analyzes the generated text data and applies an algorithm to extract key points. It extracts key points from the input text data and uses them to generate concise subtitle data that highlights the essentials. This results in a concise subtitle containing only the most important information.
[0784] Step 5:
[0785] The device senses the user's emotional state in real time using its camera and microphone. This sensor data is then used to analyze the user's emotional state using the Azure Emotion Recognition API. The analysis results are output, and the emotional state is quantified.
[0786] Step 6:
[0787] The server dynamically adjusts subtitle styling and display time based on user emotion data. Emotion data is used as input, and the color, font, and timing of subtitle display are changed in real time. This creates subtitles that visually reflect emotions.
[0788] Step 7:
[0789] The server generates adjusted subtitle data and outputs it in a format editable with Adobe Premiere Pro and other software. The output project file format is then provided to the user, allowing them to edit the video according to their own needs.
[0790] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0791] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0792] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0793] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0794] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0795] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0796] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0797] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0798] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0799] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0800] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0801] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0802] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0803] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0804] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0805] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0806] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0807] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0808] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0809] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0810] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0811] The following is further disclosed regarding the embodiments described above.
[0812] (Claim 1)
[0813] A means for extracting audio from input video data,
[0814] A means of converting speech into text data using speech recognition technology,
[0815] A method for extracting key points from text data and generating subtitles that highlight those key points,
[0816] A means for identifying speakers and displaying subtitles in different fonts or colors for each speaker,
[0817] A method for outputting subtitle data in an editable project file format,
[0818] A system that includes this.
[0819] (Claim 2)
[0820] The system according to claim 1, further comprising means for storing rules for setting different display formats based on speaker identification.
[0821] (Claim 3)
[0822] The system according to claim 1, further comprising means for providing generated subtitle data to a user and for improving the subtitle generation process based on user feedback.
[0823] "Example 1"
[0824] (Claim 1)
[0825] A method for extracting audio information from video information,
[0826] A means for converting speech information into text information using a speech recognition method,
[0827] A means for extracting important information from textual information and generating display information that highlights the important information,
[0828] A means for identifying the speaker and displaying information in a different display format for each speaker,
[0829] A means of outputting the displayed information in an editable project information format,
[0830] A system that includes this.
[0831] (Claim 2)
[0832] The system according to claim 1, further comprising means for storing rules for setting different display formats based on speaker identification.
[0833] (Claim 3)
[0834] The system according to claim 1, further comprising means for providing generated display information to a user and for improving the display information generation process based on feedback from the user.
[0835] "Application Example 1"
[0836] (Claim 1)
[0837] A means for extracting audio from input video data,
[0838] A means of converting speech into text data using speech recognition technology,
[0839] A means for extracting key points from text data and generating display information that summarizes those key points,
[0840] A means for identifying speakers and presenting display information using different character styles or colors for each speaker,
[0841] A means of outputting project data in an editable format,
[0842] A control means that generates and visually presents display information in real time,
[0843] A system that includes this.
[0844] (Claim 2)
[0845] The system according to claim 1, further comprising means for storing rules for setting different display formats based on speaker identification.
[0846] (Claim 3)
[0847] The system according to claim 1, further comprising means for providing generated display information to a user and for improving the display information generation process based on evaluations from the user.
[0848] "Example 2 of combining an emotion engine"
[0849] (Claim 1)
[0850] A means for extracting audio from input video information,
[0851] A means of converting speech into text information using speech recognition techniques,
[0852] A method for extracting key points from text information and generating subtitles that capture those key points,
[0853] A means of recognizing the user's emotions and dynamically adjusting the subtitle style based on those emotions,
[0854] A means of outputting subtitle data in an editable format,
[0855] A system that includes this.
[0856] (Claim 2)
[0857] The system according to claim 1, further comprising means for storing rules for setting display formats based on emotion recognition.
[0858] (Claim 3)
[0859] The system according to claim 1, further comprising means for providing generated subtitle data to a user and for improving the subtitle generation process based on feedback from the user.
[0860] "Application example 2 when combining with an emotional engine"
[0861] (Claim 1)
[0862] A means for extracting audio from input video data,
[0863] A means of converting speech into text data using speech recognition technology,
[0864] A method for extracting key points from text data and generating subtitles that highlight those key points,
[0865] A means for identifying speakers and displaying subtitles with different display characteristics for each speaker,
[0866] A means for dynamically adjusting the styling and display time of subtitles based on the user's emotional state,
[0867] A means of sensing emotions using an input device in order to analyze emotional states,
[0868] A means of outputting subtitle data in an editable data format,
[0869] A system that includes this.
[0870] (Claim 2)
[0871] A means for storing rules for setting different display characteristics based on speaker identification,
[0872] The system according to claim 1, further comprising means for memorizing user emotion information and learning subtitle generation based on usage history.
[0873] (Claim 3)
[0874] The system according to claim 1, further comprising means for providing generated subtitle data to a user and improving the subtitle generation process based on user feedback, and means for proposing interactions generated by analyzing visual and auditory information in real time. [Explanation of Symbols]
[0875] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A method for extracting audio from input video data, A means of converting speech into text data using speech recognition technology, A method for extracting key points from text data and generating subtitles that highlight those key points, A means for identifying speakers and displaying subtitles in different fonts or colors for each speaker, A method for outputting subtitle data in an editable project file format, A system that includes this.
2. The system according to claim 1, further comprising means for storing rules for setting different display formats based on speaker identification.
3. The system according to claim 1, further comprising means for providing generated subtitle data to a user and for improving the subtitle generation process based on user feedback.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A