System
The system efficiently generates and distributes high-quality clipped videos by analyzing, extracting highlight scenes, and adding comments, addressing inefficiencies in conventional methods.
Patent Information
- Application Number
- JP2024125416
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2026-02-13
AI Technical Summary
Conventional methods for creating high-quality clipped videos from large numbers of videos are inefficient due to the extensive effort required in analyzing, selecting highlight scenes, generating comments, converting content formats, and distributing them, making it difficult for content creators and agencies.
A system that includes uploading video files, analyzing them to extract highlight scenes, generating comments, overlaying comments on the video, converting the format, and distributing the video to specific platforms, with additional checks for metadata and keyword-based scene extraction for accuracy.
Enables efficient and accurate generation of high-quality clipped videos, allowing users to easily upload, create, and distribute these videos quickly on various platforms.
Smart Images

Figure 2026023481000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The problem that this invention aims to solve is to lower the technical and time hurdles for creating effective clips from a large number of videos in response to the increasing demand for high-quality clipped videos to attract users due to the rapid increase in video content. Conventional methods have had the problem that the entire process, from analyzing the videos to selecting highlight scenes, creating comments, converting the content format, and distributing it, requires a great deal of effort, making them inefficient for content creators and agencies. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by providing a system including a means for uploading a video file, a means for analyzing the video file, a means for extracting highlight scenes based on the analysis results, a means for generating comments for the extracted highlight scenes, a means for overlaying the generated comments on the video, a means for converting the overlaid highlight video into a format for a specific platform, and a means for distributing the converted video to the specific platform. Furthermore, by further providing a means for checking the metadata of the video file and performing format conversion as necessary, and a means for extracting highlight scenes based on specific keywords or scene changes during the video analysis process, more accurate and efficient generation of cut-out videos becomes possible.
[0006] A "video file" is a file that contains audio and video data recorded in digital format.
[0007] "Upload" refers to the act of transferring data from a user's device to a server.
[0008] "Analysis" is the process of examining data in detail and extracting specific information or features.
[0009] A "highlight scene" is a particularly important or noteworthy part of a video.
[0010] A "comment" is a written explanation or comment on a scene in a video.
[0011] "Overlay" is a technique for superimposing information or graphics onto another image or data.
[0012] "Format conversion" is the process of changing the format or structure of data into another format.
[0013] "Distribution" refers to the act of transmitting content to a particular platform or user and making it publicly available. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] The system of this invention uses AI to automatically generate highlight scenes and comments from videos, creating high-quality clipped videos. In this system, the user, server, and device each play an important role.
[0036] Program processing
[0037] Uploading videos
[0038] Users upload video files to the server using their devices (PCs, smartphones, tablets, etc.). The uploaded video files are then stored in the server's storage.
[0039] Preparing for video analysis
[0040] The server checks the metadata of the uploaded video file and converts the format as needed, making it easier for the generation AI to analyze.
[0041] Video Analysis
[0042] The server passes the prepared video file to the AI generator and requests it to analyze the video. The AI analyzes the video in detail and identifies important scenes (highlight scenes) based on specific keywords and visual changes.
[0043] Highlight Scene Extraction
[0044] Based on the analysis results, the generation AI extracts multiple highlight scenes, which constitute the parts that are particularly interesting to the user.
[0045] Comment Generation
[0046] The server automatically adds comments to the highlight scenes using AI, which are generated using natural language generation technology and include explanations and impressions of the scenes.
[0047] Generate cropped video
[0048] The server combines the extracted highlights with the generated comments to create a clipped video, during which the generated comments are added as text overlays to the video.
[0049] Format Conversion
[0050] The server then converts the resulting cropped video into the format of the specific platform (e.g., social networking site, video streaming site, etc.), which may include changing the aspect ratio and resolution of the video.
[0051] delivery
[0052] The server delivers the optimized cut-out video to a specific platform, allowing users to enjoy high-quality cut-out video on that platform.
[0053] Specific examples
[0054] For example, consider the case where a user uploads a video of a soccer match from their smartphone to a server. When the user uploads the video file of the match, the server receives the file and converts the format if necessary. The server then passes the video file to a generation AI, which detects the score of the match, the cheers of the crowd, and so on.
[0055] The generation AI extracts particularly noteworthy scenes from a match and generates comments for them, such as "Amazing goal!" or "Great play!" The server combines these highlight scenes and comments to create a clipped video that users can easily watch. The server then converts the video into a vertical format suitable for social media platforms and automatically posts it. As a result, users can quickly obtain an appealing clipped video and share it on social media.
[0056] In this way, the system of the present invention provides a powerful means for users to easily create and distribute high-quality clipped videos through advanced analysis and automatic processing of video content.
[0057] The processing flow will be explained below.
[0058] Step 1:
[0059] The user uploads a video file to the server using the device. The user opens a dedicated upload screen, selects the video file, and clicks the send button.
[0060] Step 2:
[0061] The server receives the video file and stores it in temporary storage. The server checks the metadata of the video file, recording the file size and format. The server saves the received video file in the specified directory.
[0062] Step 3:
[0063] The server checks the format of the video file and converts it if necessary. If the file format is not a standard format (e.g. MP4), a conversion tool is used to convert it to the appropriate format.
[0064] Step 4:
[0065] The server calls an API to pass the prepared video file to the generation AI and starts the analysis job. The server records the analysis progress as a log and monitors the progress while the generation AI performs the analysis.
[0066] Step 5:
[0067] The generative AI analyzes the content of the video and extracts important scenes (highlight scenes) based on specific keywords (e.g., scores and cheers in sports), scene changes, and visual emphasis.
[0068] Step 6:
[0069] The server then requests the AI to automatically generate comments for the highlight scenes obtained from the AI. The AI then uses natural language generation technology to generate comments appropriate for the highlight scenes.
[0070] Step 7:
[0071] The server combines the highlight scenes and the generated comments to create a clipped video. Using a video editing library, the server clips out the specified highlight scenes and adds the generated comments to the video as a text overlay.
[0072] Step 8:
[0073] The server converts the resulting cropped video to the format of the specific platform, converting the video aspect ratio and resolution to a platform-appropriate format (e.g., vertical, 1080x1920).
[0074] Step 9:
[0075] The server distributes the converted video to a specific platform, and posts the clipped video to a specified account using the platform's API, such as a social networking site. This allows users to obtain high-quality clipped videos in a short amount of time.
[0076] Example 1
[0077] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0078] Analyzing and editing video data takes time and effort, especially extracting highlight scenes, generating captions for those scenes, and converting them into formats suitable for specific platforms. This makes it difficult for users to quickly create and distribute high-quality clipped videos.
[0079] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0080] In this invention, the server includes means for uploading video data, means for checking metadata of the video data and converting the format as necessary, means for analyzing the video data, means for extracting important scenes based on the analysis results, means for generating explanatory text for the extracted important scenes, means for overlaying the generated explanatory text on the video, means for converting the overlaid important video into a format for a specific platform, and means for delivering the converted video to the specific platform. This allows users to easily upload video data and quickly generate and deliver high-quality cut-out videos.
[0081] "Video data" means digital files containing visual and audio information.
[0082] "Metadata" is additional information that accompanies video data, including technical details such as file format, resolution, frame rate, bit rate, etc.
[0083] "Format conversion" is the process of converting a file from one digital format to another digital format.
[0084] "Analysis" refers to the detailed investigation and analysis of digital information, and when applied to video data in particular, refers to the evaluation of the video and audio content using specialized algorithms.
[0085] An "important scene" is a particularly noteworthy portion of video data, and refers to a scene with prominent visual or audio characteristics or a scene that is likely to be of great interest to the user.
[0086] A "description" is a sentence generated for a specific scene, which expresses the content and meaning of the scene in words.
[0087] "Text overlay" is a technique for displaying text information overlaid on specific scenes in a video.
[0088] "Platform" refers to an online system or environment for distributing video data and enabling users to view it.
[0089] "Distribution" refers to the operation or process of delivering content to users, especially over the Internet.
[0090] This invention relates to a system that uses generative AI to automatically generate highlight scenes and comments from videos and create high-quality clips. By using this system, users can easily upload videos, create clips based on the analysis results, and distribute them to specific platforms.
[0091] First, a user uploads video data to the server using a device such as a PC, smartphone, or tablet. For example, a user selects a soccer game video taken with a smartphone and presses the "upload" button. The video data is then transferred to the server. A progress bar is displayed to indicate the progress.
[0092] The server then checks the metadata of the uploaded video data and converts the format if necessary. For example, video data in .avi format is converted to .mp4 format, which is easier for the generative AI model to analyze.
[0093] Once the server has completed preparing the video data, it passes the data to the generative AI and requests it to analyze it. The generative AI model is then sent a prompt such as, "Please identify the highlight scenes in this video." At this time, the generative AI model detects specific keywords (e.g., "goal" or "cheers") and visual changes in the video to identify important scenes.
[0094] The AI generator creates a list of important scenes (highlights) based on the analysis results. For example, you can set the time range and scene content, such as "00:05:20-00:05:40: The goal scene" or "00:10:10-00:10:30: The crowd cheering scene."
[0095] The server then asks the AI to generate a description for the highlight scene. For example, it sends a prompt such as, "Please generate a description for this scene." The AI generates a comment for the "goal scene" such as, "A great shot to score!"
[0096] The server then combines the extracted highlights with the generated captions to create a short video clip. For example, it adds a text overlay to a goal highlight, saying "Great shot, score!", and saves it as a short video clip.
[0097] The server then converts the finished cropped video to the format for the specific platform, for example converting it from a 16:9 landscape format for YouTube to a 9:16 portrait format for Instagram, and adjusting the resolution to 1080p.
[0098] Finally, the server distributes the optimized video to a specific platform, such as Instagram, where it automatically posts the video to the user's account, automatically setting the video title and tags.
[0099] In this way, the user, server, and device work together to easily upload video data and quickly generate and distribute high-quality clipped videos. A specific example of a prompt is as follows:
[0100] "Identify the highlights in this video and generate comments for the parts that include goals and crowd cheers."
[0101] "Extract important scenes from a soccer game video taken with a smartphone and generate a text overlay for them."
[0102] "Analyze the most exciting moments in the game, add engaging commentary and create clips."
[0103] This allows users to easily upload video data and quickly generate and distribute high-quality clipped videos.
[0104] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0105] Step 1:
[0106] Users upload video data to a server using devices such as PCs, smartphones, and tablets. For example, when a user selects a video they have taken with their smartphone and presses the "upload" button, the video data is transferred to the server. The input is the video data uploaded from the user's device, and the output is the video data saved on the server.
[0107] Step 2:
[0108] The server checks the metadata of the uploaded video data and determines whether format conversion is necessary. For example, the server analyzes the video file format, resolution, frame rate, etc., and converts it to .mp4 format if necessary. The input is the original video data stored on the server, and the output is video data converted into a format that is easy for the generative AI to analyze.
[0109] Step 3:
[0110] The server passes the prepared video data to the generation AI and requests analysis using a prompt. For example, the prompt "Please identify the highlight scene in this video" is sent to the generation AI. The generation AI analyzes the video data and detects specific keywords and visual changes. The input is the video data stored on the server and the prompt, and the output is a list of candidate highlight scenes as the analysis result.
[0111] Step 4:
[0112] The generation AI lists important scenes based on the analysis results. For example, you can set the time range and scene content, such as "00:05:20-00:05:40 goal scene" or "00:10:10-00:10:30 crowd cheering scene." The input is the analysis result data, and the output is a list containing detailed time ranges and scene content.
[0113] Step 5:
[0114] The server requests the generation AI to generate a description for each highlight scene. For example, it sends a prompt saying, "Please generate a description for this scene." The generation AI generates an appropriate description for that scene. The input is the highlight scene data and the prompt, and the output is a description corresponding to each scene.
[0115] Step 6:
[0116] The server combines the extracted highlight scenes with the generated explanatory text as a text overlay to create a clipped video. For example, a text overlay saying "Great shot and score!" is added to a goal scene. The input is the video data of the highlight scenes and explanatory text, and the output is an edited clipped video.
[0117] Step 7:
[0118] The server then converts the completed cropped video to a format suitable for the specific platform, e.g., converting from 16:9 to 9:16 portrait format for YouTube or Instagram. The input is the cropped video, and the output is the converted video for the specific platform.
[0119] Step 8:
[0120] The server delivers the optimized video to a specific platform, for example, using the Instagram API to automatically post the video to a user's account. The input is the platform-optimized video data, and the output is the delivered video.
[0121] In this way, through each processing step of the system, users can easily generate high-quality cutout videos and share them on specific platforms.
[0122] (Application example 1)
[0123] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0124] Conventional video distribution systems lack the technology to automatically identify and extract important scenes that users may have missed while watching, and notify the viewing device of the missed scenes at the appropriate time. This causes users to miss important scenes, resulting in a poor viewing experience. Furthermore, there is a lack of a means to immediately provide appropriate explanations for extracted important scenes, making it difficult for viewers to understand the significance of the scenes.
[0125] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0126] In this invention, the server includes means for uploading video data, means for analyzing the video data, means for extracting important scenes based on the analysis results, means for generating explanations for the extracted important scenes, means for overlaying the generated explanations on the video, means for converting the overlaid important scene video into a specific distribution format, means for transmitting the converted video to a specific distribution platform, means for automatically extracting important scenes from the video being viewed in real time and immediately notifying and displaying them on the user's display device, and means for generating explanations for the important scenes automatically extracted from the video being viewed in real time and automatically providing the explanations using a generation AI, thereby enabling users to immediately understand the significance of important scenes without missing them while viewing.
[0127] "Video data" refers to digital files of moving images that have been shot or created.
[0128] A "server" refers to a computer system that provides content and services in a network environment.
[0129] "Upload" refers to the operation of sending and saving data from a user's device to a server.
[0130] "Analysis" refers to the process of examining and evaluating uploaded video data to recognize specific elements or scenes.
[0131] An "important scene" refers to a scene or event in video data that should be of particular interest.
[0132] "Extraction" refers to the process of extracting specific scenes or information based on the analysis results.
[0133] "Description" refers to information that includes commentary and impressions about a particular scene.
[0134] "Overlay" refers to the process of displaying the generated explanation on a specific scene of the original video data.
[0135] "Delivery format" refers to the video specifications and formats supported by a particular platform.
[0136] "Transmission" refers to the act of delivering the converted video to a specific platform via a network.
[0137] "Real-time" refers to processing immediately at the time the video is being captured or streamed.
[0138] "Display device" refers to a hardware device that enables a user to visually receive information.
[0139] "Notification" refers to an operation that notifies a user of specific information or events.
[0140] "Generative AI" refers to artificial intelligence technology that automatically creates text and audio data based on visual and audio information.
[0141] The system of this invention uses generation AI to automatically generate important scenes and explanations from videos, creating high-quality clipped videos. In this system, the user, server, and device each play an important role.
[0142] Program Description:
[0143] Hardware and software used:
[0144] Hardware
[0145] Display devices (e.g., smart glasses, head-mounted displays)
[0146] Server (high performance computer system)
[0147] software
[0148] OpenCV (library for image analysis)
[0149] PyTorch (an execution environment for generative AI models)
[0150] Transformers library (models for natural language generation)
[0151] System behavior:
[0152] First, the user uploads the video data they have shot or created to the server via their device. The server then checks the meta information of the received video data and converts the video format as necessary. This process prepares the video data in a format that is easy for the generative AI model to analyze.
[0153] The server then analyzes the video data to detect specific words and scene changes. Based on this analysis, the system extracts important scenes from the video, such as a goal or the cheers of the crowd in a sports game.
[0154] For the extracted important scenes, a generative AI model is used to automatically generate explanations. The generated explanations are overlaid on the video data in a visually easy-to-understand format. This overlay process allows users to instantly understand the significance of important scenes.
[0155] The overlaid video of key moments is then converted into a specified distribution format. This distribution format is adjusted to meet the specifications required by specific platforms such as social media and video distribution sites. Finally, the converted video is sent to the specific distribution platform via a server.
[0156] Examples:
[0157] For example, consider a case where a user is watching a live stream of a soccer match with smart glasses. Any goal that the user missed will be automatically detected and extracted, and a description such as "Great goal!" will be automatically generated by the generative AI model. This description will be superimposed on the display of the smart glasses.
[0158] This allows users to instantly check important scenes they missed in real time and understand their significance.
[0159] Example prompt sentence:
[0160] "How can I automatically extract goal scenes from the live stream I'm watching and notify the user in a timely manner?"
[0161] In this way, the system of the present invention provides a powerful means for ensuring that users do not miss important scenes in a video in real time and for providing instant commentary on the scenes.
[0162] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0163] Step 1:
[0164] Uploading video data
[0165] A user uses a terminal to upload video data that has been shot or created to a server.
[0166] Input: Locally stored video file
[0167] Output: Video data stored on the server
[0168] Specific operation: The user selects video data through the application interface and uploads the data by pressing the send button. The server saves the received data in storage.
[0169] Step 2:
[0170] Check and convert meta information
[0171] The server checks the meta information (e.g., resolution, frame rate, format) of the received video data and performs format conversion if necessary.
[0172] Input: Video data stored on the server
[0173] Output: Video data converted for analysis
[0174] Specific operation: The server analyzes the video metadata and converts the resolution and format as necessary to make it suitable for analysis. For example, converting an AVI file to MP4 format.
[0175] Step 3:
[0176] Video data analysis
[0177] The server passes the video data prepared for analysis to the generative AI model and performs video analysis.
[0178] Input: Format converted video data
[0179] Output: Identification information of important scenes as analysis results
[0180] Specific operation: The server uses OpenCV and PyTorch to analyze each video frame, and the generative AI model recognizes specific scenes (e.g., goal scenes).
[0181] Step 4:
[0182] Extraction of important scenes
[0183] Based on the analysis results, the server extracts important scenes from the video.
[0184] Input: Analysis results from a generative AI model
[0185] Output: Extracted frames and timestamps of key scenes
[0186] Specific operation: The server identifies important frames such as "goal scenes" and "cheering from the crowd" from the analysis results of the generative AI model and extracts their timestamp information.
[0187] Step 5:
[0188] Generate Description
[0189] The server automatically generates explanations for the extracted important scenes using a generative AI model.
[0190] Input: Extracted important scenes
[0191] Output: Generated description text
[0192] What it does: The server uses the Transformers library to generate a description for each key moment, such as "Amazing goal!" or "Great play!"
[0193] Step 6:
[0194] Description overlay
[0195] The server overlays the generated explanation onto the video data.
[0196] Input: frames of key scenes and generated explanatory text
[0197] Output: Video of important scenes with overlaid explanations
[0198] Specific operation: The server uses OpenCV to overlay the generated explanatory text on specific frames and process it into a visually easy-to-understand form.
[0199] Step 7:
[0200] Conversion to delivery format
[0201] The server converts the superimposed important scene video into a specified distribution format.
[0202] Input: Video of key scenes with overlaid descriptions
[0203] Output: Video file converted to distribution format
[0204] Specific operation: The server adjusts the video resolution, aspect ratio, codec, etc. according to the platform specifications and converts it to the appropriate format.
[0205] Step 8:
[0206] Sending videos
[0207] The server sends the converted video file to a specific distribution platform.
[0208] Input: Video files converted to a distribution format
[0209] Output: Video published on a distribution platform
[0210] Specific operation: The server uploads the video through the specified API or endpoint and automatically publishes it.
[0211] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0212] The system of this invention uses a generative AI and an emotion engine to automatically generate highlight scenes and comments from videos and create high-quality clipped videos. In this system, the user, server, and emotion engine each play an important role.
[0213] Program processing
[0214] Uploading videos
[0215] Users upload video files to the server using their devices (PCs, smartphones, tablets, etc.). The uploaded video files are then stored in the server's storage.
[0216] Preparing for video analysis
[0217] The server checks the metadata of the uploaded video file and converts the format as needed, making it easier for the generation AI to analyze.
[0218] Video Analysis
[0219] The server passes the prepared video file to the AI generator and requests it to analyze the video. The AI analyzes the video in detail and identifies important scenes (highlight scenes) based on specific keywords and visual changes.
[0220] Highlight Scene Extraction and User Emotion Recognition
[0221] The generation AI extracts multiple highlight scenes based on the analysis results. At the same time, the emotion engine recognizes and collects the user's emotions while watching. The user's emotional data is reflected in the selection of highlight scenes.
[0222] Comment Generation
[0223] The server takes into account the user's emotional data obtained from the emotion engine and adds comments automatically generated by the AI to the highlight scenes. The comments are generated using natural language generation technology, including explanations and impressions of the scenes.
[0224] Generate cropped video
[0225] The server combines the extracted highlights with the generated comments to create a clipped video, during which the generated comments are added as text overlays to the video.
[0226] Format Conversion
[0227] The server then converts the resulting cropped video into the format of the specific platform (e.g., social networking site, video streaming site, etc.), which may include changing the aspect ratio and resolution of the video.
[0228] delivery
[0229] The server then distributes the optimized clipped video to a specific platform, allowing users to enjoy high-quality clipped videos on that platform. The distributed video also reflects the user's emotional information, making it more likely to resonate with viewers.
[0230] Specific examples
[0231] For example, consider the case where a user uploads a video of a soccer match from their smartphone to a server. When the user uploads the video file of the match, the server receives the file and converts the format if necessary. The server then passes the video file to a generation AI, which detects the score of the match, the cheers of the crowd, and so on.
[0232] At the same time, the emotion engine recognizes the user's emotions in real time as they watch and collects them as data. For example, if a user becomes excited at a goal, that emotional data is reflected in the generation AI. The generation AI then extracts particularly noteworthy scenes from the game. The server then generates comments such as "Amazing goal!" or "Great play!" Comments and highlight selection that reflect the emotional data evoke even stronger empathy in viewers.
[0233] The server combines these highlights and comments to create a clipped video that users can easily watch.The server then converts the video into a vertical format suitable for social media platforms and automatically posts it.As a result, users can quickly obtain high-quality clipped videos that reflect emotional information and share them on social media.
[0234] In this way, the system of the present invention utilizes advanced analysis of video content as well as user emotional data to automatically generate and efficiently distribute more attractive and relatable clipped videos.
[0235] The processing flow will be explained below.
[0236] Step 1:
[0237] The user uploads a video file to the server using the device. The user opens a dedicated upload screen, selects the video file, and clicks the send button.
[0238] Step 2:
[0239] The server receives the video file and stores it in temporary storage. The server checks the metadata of the video file, recording the file size and format. The server saves the received video file in the specified directory.
[0240] Step 3:
[0241] The server checks the format of the video file and converts it if necessary. If the file format is not a standard format (e.g. MP4), a conversion tool is used to convert it to the appropriate format.
[0242] Step 4:
[0243] The server calls an API to pass the prepared video file to the generation AI and starts the analysis job. The server records the analysis progress as a log and monitors the progress while the generation AI performs the analysis.
[0244] Step 5:
[0245] The generative AI analyzes the content of the video and extracts important scenes (highlight scenes) based on specific keywords (e.g., scores and cheers in sports), scene changes, and visual emphasis.
[0246] Step 6:
[0247] The emotion engine recognizes the user's emotions in real time while watching a video and collects them as data. User emotion data is acquired from facial expressions, voice, input information, etc. while watching a video.
[0248] Step 7:
[0249] The server again requests the AI to automatically generate comments for the highlight scenes obtained from the AI, taking into account the user's emotional data obtained from the emotion engine.
[0250] Step 8:
[0251] The server combines the highlight scenes and the generated comments to create a clipped video. Using a video editing library, the server clips out the specified highlight scenes and adds the generated comments to the video as a text overlay.
[0252] Step 9:
[0253] The server converts the resulting cropped video to the format of the specific platform, converting the video aspect ratio and resolution to a platform-appropriate format (e.g., vertical, 1080x1920).
[0254] Step 10:
[0255] The server distributes the converted video to a specific platform, and posts the clipped video to a specified account using the platform's API, such as a social networking site. This allows users to quickly obtain high-quality clipped videos that reflect their emotions.
[0256] Example 2
[0257] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0258] Conventional video editing systems require users to manually extract highlights and add comments, which requires time and effort. It is also difficult to achieve sophisticated editing that reflects the user's emotions, making it difficult to elicit empathy from viewers. Furthermore, format conversion and distribution for different platforms must be done manually, resulting in inefficiency.
[0259] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for uploading video files, means for analyzing video files, means for extracting highlight scenes based on the analysis results, means for generating comments for the extracted highlight scenes, means for recognizing and collecting user emotions, means for generating comments taking into account the collected emotion data, means for overlaying the generated comments on a video, means for converting the overlaid highlight video into a format for a specific platform, and means for delivering the converted video to a specific platform. This makes it possible to automatically generate high-quality clipped videos that reflect user emotions and efficiently deliver them to multiple platforms.
[0260] "Means for uploading video files" refers to the function that allows users to use their devices to send and save video files to a server.
[0261] The "means for analyzing video files" is a function that allows the server to automatically analyze the contents of uploaded video files.
[0262] The "means for extracting highlight scenes based on the analysis results" is a function for extracting specific important scenes based on the content analysis of the video.
[0263] The "means for generating a comment for the extracted highlight scene" is a function for automatically generating a comment in accordance with the extracted highlight scene.
[0264] "Means for recognizing and collecting user emotions" is a function that detects the user's emotions while watching in real time and collects them as data.
[0265] The "means for generating comments taking into consideration collected emotional data" is a function for generating comments that evoke greater empathy based on the user's emotional data.
[0266] The "means for overlaying generated comments on video" is a function for displaying generated comments overlaid on highlight scenes.
[0267] The "means for converting the overlaid highlight video into a format for a specific platform" is a function for converting the edited video into a format suitable for the standards of each platform.
[0268] "Means for distributing the converted video to a specific platform" refers to the function of uploading and distributing the format-converted video to platforms such as social media and video distribution sites.
[0269] The following describes an embodiment of the present invention. The system of the present invention utilizes a generation AI and an emotion engine to automatically generate highlight scenes and comments from videos and create high-quality clipped videos.
[0270] Uploading a video file
[0271] Users use their devices (PCs, smartphones, tablets, etc.) to upload video files to the server. They access the system's web application, log in, select the video file, and click the upload button. This video file is then saved in the server's storage. Specifically, a cloud storage service (e.g., Amazon S3) may be used.
[0272] Preparing for video analysis
[0273] The server receives the uploaded video file and checks its metadata. If necessary, it converts the video format to make it easier for the AI to analyze. For example, it uses software called FFmpeg to convert the original video format to H.264.
[0274] Video Analysis
[0275] The server passes the prepared video file to the generation AI, which uses, for example, the OpenAI API. The server inputs a prompt to the generation AI, such as "I want you to detect soccer goal scenes." The generation AI then performs a detailed analysis of specific keywords and visual changes in the video, outputting the timestamp and characteristics of each scene.
[0276] Highlight Scene Extraction and User Emotion Recognition
[0277] The generation AI extracts multiple highlight scenes based on the analysis results. At the same time, the emotion engine recognizes and collects the user's emotions in real time as they watch. This emotion recognition is performed using tools such as EmoPy. As the user watches the video, the server reads the user's facial expressions through the device's webcam. For example, if the user smiles during a goal-scoring scene, that emotional data is recorded.
[0278] Comment Generation
[0279] The server passes the user's emotional data obtained from the emotion engine to the generation AI. The generation AI generates comments such as "Incredible goal!" or "Great play!" based on data such as "This scene made the user smile." Specifically, natural language generation technology is often used. For example, the GPT-4 model is used.
[0280] Generate cropped video
[0281] The server combines the extracted highlights with the generated comments to create a clipped video, which is then added as a text overlay to the video. The video is edited using software such as Adobe Premiere Pro.
[0282] Format Conversion
[0283] The server then converts the resulting cut video into the format of the specific platform (e.g., social media, video streaming site, etc.), using tools such as FFmpeg to change the aspect ratio and resolution of the video.
[0284] delivery
[0285] The server automatically posts the optimized clip to a specific platform, often via the YouTube API or a social networking service API. Users receive a notification, watch the video, and share it on social media.
[0286] In this way, it is possible to automatically generate high-quality clipped videos that reflect the user's emotions and efficiently distribute them across multiple platforms, significantly reducing the time-consuming and labor-intensive work that was previously done manually, and providing video content that evokes empathy among viewers.
[0287] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0288] Step 1:
[0289] A user accesses the system's web application using a device (PC, smartphone, tablet, etc.) and logs in. Next, the user selects a video file and clicks the upload button. The input is the video file selected by the user, and the output is the video data sent to the server. The server stores this file in cloud storage (e.g., Amazon S3). The specific operation is an HTTP POST request.
[0290] Step 2:
[0291] The server receives the uploaded video file and checks the video file's metadata (resolution, format, length, etc.). If necessary, it uses FFmpeg to convert the original video format to H.264 format. The input is the original uploaded video file, and the output is the format-converted video file. Specific operations involve the execution of FFmpeg commands.
[0292] Step 3:
[0293] The server passes the converted video file to the generation AI (e.g., OpenAI's API). The input is the format-converted video file and a prompt statement such as, "I want you to detect soccer goal scenes." The output is the analysis results (timestamp and features of the goal scene) returned by the generation AI. Specific operations include sending an API request and a response.
[0294] Step 4:
[0295] The generation AI extracts multiple highlight scenes based on the analysis results. The server receives these and uses an emotion engine (e.g., EmoPy) to recognize and collect the user's emotions in real time while watching. The input is the analysis results and the user's real-time video (facial expression data), and the output is a highlight scene list including the user's emotional data. Specific operation uses a data stream via WebSocket.
[0296] Step 5:
[0297] The server passes the emotion data obtained from the emotion engine to the generation AI. Based on data such as "This scene made the user smile," the generation AI automatically generates comments such as "Incredible goal!" or "Great play!" The input is the user's emotion data and a list of highlight scenes, and the output is the generated comments. Specific operations include sending and responding to API requests.
[0298] Step 6:
[0299] The server combines the extracted highlight scenes and the generated comments to create a cut-out video. This process uses Adobe Premiere Pro scripts or similar video editing software. The input is the highlight scene list and the generated comments, and the output is the completed cut-out video. Specific operations include timeline editing and command execution.
[0300] Step 7:
[0301] The server uses FFmpeg to convert the completed cut video into the format of a specific platform (e.g., social networking site, video distribution site). The input is the edited cut video, and the output is a format-converted video file. Specifically, the FFmpeg command is executed again.
[0302] Step 8:
[0303] The server automatically posts clipped videos optimized for specific platforms via API. The input is the format-converted video file and the platform information for posting, and the output is the URL of the uploaded video. Specifically, a request is made to the API endpoint of each platform. The user receives this URL notification, watches it, and shares it with many people.
[0304] (Application example 2)
[0305] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0306] Conventional video analysis systems simply extract highlight scenes based on specific keywords or visual changes, without considering the user's emotions while watching. This resulted in the generation of videos that failed to appeal to the user's emotions and did not evoke empathy. Furthermore, the generation of comments was also uniform, which was an issue that did not fully reflect the user's emotions.
[0307] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0308] In this invention, the server includes means for uploading video files, means for analyzing video files, means for extracting highlight scenes based on the analysis results, means for generating comments for the extracted highlight scenes, means for overlaying the generated comments on the video, means for converting the overlaid highlight video into a format for a specific platform, means for distributing the converted video to a specific content distribution service, means for sensing a user's emotions while watching in real time and collecting emotion data, and means for generating comments based on the emotion data. This makes it possible to generate highlight videos and comments that reflect the user's emotions while watching and that evoke greater empathy.
[0309] The "means for uploading video files" is a function that allows a user to send video files from their own device to the server and store them in the server's storage.
[0310] "Means for analyzing video files" refers to technology that examines the video in detail and extracts necessary information in order to understand the content of the uploaded video file.
[0311] "Means for extracting highlight scenes" refers to a technology that identifies important or noteworthy scenes as a result of video analysis and extracts them.
[0312] The "means for generating comments" is a function that uses a generative AI model to create natural language sentences for extracted highlight scenes.
[0313] "Video overlay means" refers to a technique for displaying generated comments or other information overlaid on video frames.
[0314] The "means for converting into a format" refers to a technology for changing the resolution, aspect ratio, etc. of the generated highlight video in order to optimally display it on a specific platform.
[0315] "Means of distribution" refers to the function of automatically sending and publishing the converted video to a specified content distribution service or social networking site.
[0316] "Means for sensing emotions in real time" refers to technology that detects the emotional state of a user in real time from facial expressions, voice, etc. while the user is watching a video.
[0317] The "means for collecting emotion data" refers to a technology for recording detected emotion information of a user as data and using it for subsequent analysis and processing.
[0318] "Means for generating comments based on emotional data" refers to a technology in which a generative AI model automatically creates comments that are appropriate to the user's emotions, taking into account collected emotional data.
[0319] The system for implementing the present invention includes a terminal, a server, a generative AI model, and an emotion engine. The specific operation of this system will be described in detail below.
[0320] The terminal is a device for users to upload video files, and can be a smartphone, tablet, or PC. Using this terminal, users upload video files they have taken. The video files are sent to the server and stored in the server's storage.
[0321] The server analyzes the uploaded video files and extracts important scenes (highlight scenes) based on the analysis results. This analysis uses a generative AI model to analyze the video content in detail. The server also converts the format of the video file as needed to make it easier to analyze.
[0322] Next, the server uses an emotion engine to sense the user's emotions in real time while watching the video and collects emotion data. This emotion data is acquired while the user is watching the video, and the user's emotional state is detected from their facial expressions, voice, gestures, etc. The data obtained by the emotion engine is sent to a generative AI model and reflected in the selection of highlight scenes and comment generation.
[0323] The generative AI model takes into account the collected emotional data and automatically generates appropriate comments for the extracted highlight scenes. These comments are in line with the video content and the user's emotions, aiming to evoke empathy in the viewer. The generated comments are added as a text overlay to the video.
[0324] The overlaid highlight footage is then converted into a format suitable for the specific platform (e.g. YouTube, Instagram, etc.) The server performs this conversion, adjusting the video resolution, aspect ratio, etc.
[0325] Finally, the server automatically distributes the converted highlight video to the specified content distribution service or social networking site, allowing users to easily publish and share the video.
[0326] As a concrete example, consider the case where a user uploads a video of a soccer match from their smartphone to a server. When the user uploads the video file of the match, the server receives the file and converts the format as necessary. The server then passes the video file to a generative AI model, which detects things like goal scenes and crowd cheers. At the same time, as the user watches, the emotion engine recognizes the user's emotions in real time and collects them as data. For example, if the user becomes excited at a goal, that emotional data is reflected in the generative AI model, which selects the scene as particularly noteworthy. The server then generates comments such as "Amazing goal!" or "Great play!"
[0327] Below are some example prompts to input to a generative AI model:
[0328] "Please extract the goal scenes and excitement of the crowd from this soccer match video, generate optimal comments based on the user's emotional data, and create a cut-out video."
[0329] In this way, the system can enhance the user's viewing experience and generate high-quality videos that evoke stronger empathy in the viewer.
[0330] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0331] Step 1:
[0332] The server receives video files uploaded by users from their devices and stores them in the server's storage. The input is the video file from the user's device, and the output is the video file stored in the server storage. In this step, file reception and storage operations are performed.
[0333] Step 2:
[0334] The server checks the metadata of the stored video file and performs format conversion if necessary. The input is the stored video file, and the output is a video file converted into a format that is easy for the generative AI model to analyze. In this step, metadata is checked and format conversion is performed.
[0335] Step 3:
[0336] The server passes the converted video file to the generative AI model, which then performs a detailed analysis of the video content. The input is the converted video file, and the output is the highlight scenes identified as a result of the video analysis. In this step, detailed video analysis is performed using the generative AI model.
[0337] Step 4:
[0338] The server uses an emotion engine to sense the user's emotions in real time while the user is watching and collects emotion data. The input is the user's facial expressions and voice while watching, and the output is the emotion data collected in real time. In this step, emotion recognition and data collection take place.
[0339] Step 5:
[0340] The server uses a generative AI model based on the collected emotion data to generate comments for the highlight scenes. The input is emotion data and the highlight scenes, and the output is the generated comments. In this step, emotion data is analyzed and comments are generated.
[0341] Step 6:
[0342] The server adds the generated comments to the video as a text overlay. The input is the highlight scene and the generated comments, and the output is the highlight video with the comments overlayed. In this step, the text overlay is added.
[0343] Step 7:
[0344] The server converts the overlaid highlight video into a specific platform format. The input is the highlight video with overlaid comments, and the output is a video format optimized for the specific platform. In this step, the format conversion takes place.
[0345] Step 8:
[0346] The server automatically distributes the converted highlight video to the specified content distribution service or social networking site. The input is a video file converted for a specific platform, and the output is a distributed and published highlight video. In this step, the video is distributed and published.
[0347] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0348] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0349] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0350] [Second embodiment]
[0351] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0352] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0353] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0354] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0355] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0356] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0357] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0358] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0359] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0360] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0361] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0362] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0363] The system of this invention uses AI to automatically generate highlight scenes and comments from videos, creating high-quality clipped videos. In this system, the user, server, and device each play an important role.
[0364] Program processing
[0365] Uploading videos
[0366] Users upload video files to the server using their devices (PCs, smartphones, tablets, etc.). The uploaded video files are then stored in the server's storage.
[0367] Preparing for video analysis
[0368] The server checks the metadata of the uploaded video file and converts the format as needed, making it easier for the generation AI to analyze.
[0369] Video Analysis
[0370] The server passes the prepared video file to the AI generator and requests it to analyze the video. The AI analyzes the video in detail and identifies important scenes (highlight scenes) based on specific keywords and visual changes.
[0371] Highlight Scene Extraction
[0372] Based on the analysis results, the generation AI extracts multiple highlight scenes, which constitute the parts that are particularly interesting to the user.
[0373] Comment Generation
[0374] The server automatically adds comments to the highlight scenes using AI, which are generated using natural language generation technology and include explanations and impressions of the scenes.
[0375] Generate cropped video
[0376] The server combines the extracted highlights with the generated comments to create a clipped video, during which the generated comments are added as text overlays to the video.
[0377] Format Conversion
[0378] The server then converts the resulting cropped video into the format of the specific platform (e.g., social networking site, video streaming site, etc.), which may include changing the aspect ratio and resolution of the video.
[0379] delivery
[0380] The server delivers the optimized cut-out video to a specific platform, allowing users to enjoy high-quality cut-out video on that platform.
[0381] Specific examples
[0382] For example, consider the case where a user uploads a video of a soccer match from their smartphone to a server. When the user uploads the video file of the match, the server receives the file and converts the format if necessary. The server then passes the video file to a generation AI, which detects the score of the match, the cheers of the crowd, and so on.
[0383] The generation AI extracts particularly noteworthy scenes from a match and generates comments for them, such as "Amazing goal!" or "Great play!" The server combines these highlight scenes and comments to create a clipped video that users can easily watch. The server then converts the video into a vertical format suitable for social media platforms and automatically posts it. As a result, users can quickly obtain an appealing clipped video and share it on social media.
[0384] In this way, the system of the present invention provides a powerful means for users to easily create and distribute high-quality clipped videos through advanced analysis and automatic processing of video content.
[0385] The processing flow will be explained below.
[0386] Step 1:
[0387] The user uploads a video file to the server using the device. The user opens a dedicated upload screen, selects the video file, and clicks the send button.
[0388] Step 2:
[0389] The server receives the video file and stores it in temporary storage. The server checks the metadata of the video file, recording the file size and format. The server saves the received video file in the specified directory.
[0390] Step 3:
[0391] The server checks the format of the video file and converts it if necessary. If the file format is not a standard format (e.g. MP4), a conversion tool is used to convert it to the appropriate format.
[0392] Step 4:
[0393] The server calls an API to pass the prepared video file to the generation AI and starts the analysis job. The server records the analysis progress as a log and monitors the progress while the generation AI performs the analysis.
[0394] Step 5:
[0395] The generative AI analyzes the content of the video and extracts important scenes (highlight scenes) based on specific keywords (e.g., scores and cheers in sports), scene changes, and visual emphasis.
[0396] Step 6:
[0397] The server then requests the AI to automatically generate comments for the highlight scenes obtained from the AI. The AI then uses natural language generation technology to generate comments appropriate for the highlight scenes.
[0398] Step 7:
[0399] The server combines the highlight scenes and the generated comments to create a clipped video. Using a video editing library, the server clips out the specified highlight scenes and adds the generated comments to the video as a text overlay.
[0400] Step 8:
[0401] The server converts the resulting cropped video to the format of the specific platform, converting the video aspect ratio and resolution to a platform-appropriate format (e.g., vertical, 1080x1920).
[0402] Step 9:
[0403] The server distributes the converted video to a specific platform, and posts the clipped video to a specified account using the platform's API, such as a social networking site. This allows users to obtain high-quality clipped videos in a short amount of time.
[0404] Example 1
[0405] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0406] Analyzing and editing video data takes time and effort, especially extracting highlight scenes, generating captions for those scenes, and converting them into formats suitable for specific platforms. This makes it difficult for users to quickly create and distribute high-quality clipped videos.
[0407] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0408] In this invention, the server includes means for uploading video data, means for checking metadata of the video data and converting the format as necessary, means for analyzing the video data, means for extracting important scenes based on the analysis results, means for generating explanatory text for the extracted important scenes, means for overlaying the generated explanatory text on the video, means for converting the overlaid important video into a format for a specific platform, and means for delivering the converted video to the specific platform. This allows users to easily upload video data and quickly generate and deliver high-quality cut-out videos.
[0409] "Video data" means digital files containing visual and audio information.
[0410] "Metadata" is additional information that accompanies video data, including technical details such as file format, resolution, frame rate, bit rate, etc.
[0411] "Format conversion" is the process of converting a file from one digital format to another digital format.
[0412] "Analysis" refers to the detailed investigation and analysis of digital information, and when applied to video data in particular, refers to the evaluation of the video and audio content using specialized algorithms.
[0413] An "important scene" is a particularly noteworthy portion of video data, and refers to a scene with prominent visual or audio characteristics or a scene that is likely to be of great interest to the user.
[0414] A "description" is a sentence generated for a specific scene, which expresses the content and meaning of the scene in words.
[0415] "Text overlay" is a technique for displaying text information overlaid on specific scenes in a video.
[0416] "Platform" refers to an online system or environment for distributing video data and enabling users to view it.
[0417] "Distribution" refers to the operation or process of delivering content to users, especially over the Internet.
[0418] This invention relates to a system that uses generative AI to automatically generate highlight scenes and comments from videos and create high-quality clips. By using this system, users can easily upload videos, create clips based on the analysis results, and distribute them to specific platforms.
[0419] First, a user uploads video data to the server using a device such as a PC, smartphone, or tablet. For example, a user selects a soccer game video taken with a smartphone and presses the "upload" button. The video data is then transferred to the server. A progress bar is displayed to indicate the progress.
[0420] The server then checks the metadata of the uploaded video data and converts the format if necessary. For example, video data in .avi format is converted to .mp4 format, which is easier for the generative AI model to analyze.
[0421] Once the server has completed preparing the video data, it passes the data to the generative AI and requests it to analyze it. The generative AI model is then sent a prompt such as, "Please identify the highlight scenes in this video." At this time, the generative AI model detects specific keywords (e.g., "goal" or "cheers") and visual changes in the video to identify important scenes.
[0422] The AI generator creates a list of important scenes (highlights) based on the analysis results. For example, you can set the time range and scene content, such as "00:05:20-00:05:40: The goal scene" or "00:10:10-00:10:30: The crowd cheering scene."
[0423] The server then asks the AI to generate a description for the highlight scene. For example, it sends a prompt such as, "Please generate a description for this scene." The AI generates a comment for the "goal scene" such as, "A great shot to score!"
[0424] The server then combines the extracted highlights with the generated captions to create a short video clip. For example, it adds a text overlay to a goal highlight, saying "Great shot, score!", and saves it as a short video clip.
[0425] The server then converts the finished cropped video to the format for the specific platform, for example converting it from a 16:9 landscape format for YouTube to a 9:16 portrait format for Instagram, and adjusting the resolution to 1080p.
[0426] Finally, the server distributes the optimized video to a specific platform, such as Instagram, where it automatically posts the video to the user's account, automatically setting the video title and tags.
[0427] In this way, the user, server, and device work together to easily upload video data and quickly generate and distribute high-quality clipped videos. A specific example of a prompt is as follows:
[0428] "Identify the highlights in this video and generate comments for the parts that include goals and crowd cheers."
[0429] "Extract important scenes from a soccer game video taken with a smartphone and generate a text overlay for them."
[0430] "Analyze the most exciting moments in the game, add engaging commentary and create clips."
[0431] This allows users to easily upload video data and quickly generate and distribute high-quality clipped videos.
[0432] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0433] Step 1:
[0434] Users upload video data to a server using devices such as PCs, smartphones, and tablets. For example, when a user selects a video they have taken with their smartphone and presses the "upload" button, the video data is transferred to the server. The input is the video data uploaded from the user's device, and the output is the video data saved on the server.
[0435] Step 2:
[0436] The server checks the metadata of the uploaded video data and determines whether format conversion is necessary. For example, the server analyzes the video file format, resolution, frame rate, etc., and converts it to .mp4 format if necessary. The input is the original video data stored on the server, and the output is video data converted into a format that is easy for the generative AI to analyze.
[0437] Step 3:
[0438] The server passes the prepared video data to the generation AI and requests analysis using a prompt. For example, the prompt "Please identify the highlight scene in this video" is sent to the generation AI. The generation AI analyzes the video data and detects specific keywords and visual changes. The input is the video data stored on the server and the prompt, and the output is a list of candidate highlight scenes as the analysis result.
[0439] Step 4:
[0440] The generation AI lists important scenes based on the analysis results. For example, you can set the time range and scene content, such as "00:05:20-00:05:40 goal scene" or "00:10:10-00:10:30 crowd cheering scene." The input is the analysis result data, and the output is a list containing detailed time ranges and scene content.
[0441] Step 5:
[0442] The server requests the generation AI to generate a description for each highlight scene. For example, it sends a prompt saying, "Please generate a description for this scene." The generation AI generates an appropriate description for that scene. The input is the highlight scene data and the prompt, and the output is a description corresponding to each scene.
[0443] Step 6:
[0444] The server combines the extracted highlight scenes with the generated explanatory text as a text overlay to create a clipped video. For example, a text overlay saying "Great shot and score!" is added to a goal scene. The input is the video data of the highlight scenes and explanatory text, and the output is an edited clipped video.
[0445] Step 7:
[0446] The server then converts the completed cropped video to a format suitable for the specific platform, e.g., converting from 16:9 to 9:16 portrait format for YouTube or Instagram. The input is the cropped video, and the output is the converted video for the specific platform.
[0447] Step 8:
[0448] The server delivers the optimized video to a specific platform, for example, using the Instagram API to automatically post the video to a user's account. The input is the platform-optimized video data, and the output is the delivered video.
[0449] In this way, through each processing step of the system, users can easily generate high-quality cutout videos and share them on specific platforms.
[0450] (Application example 1)
[0451] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0452] Conventional video distribution systems lack the technology to automatically identify and extract important scenes that users may have missed while watching, and notify the viewing device of the missed scenes at the appropriate time. This causes users to miss important scenes, resulting in a poor viewing experience. Furthermore, there is a lack of a means to immediately provide appropriate explanations for extracted important scenes, making it difficult for viewers to understand the significance of the scenes.
[0453] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0454] In this invention, the server includes means for uploading video data, means for analyzing the video data, means for extracting important scenes based on the analysis results, means for generating explanations for the extracted important scenes, means for overlaying the generated explanations on the video, means for converting the overlaid important scene video into a specific distribution format, means for transmitting the converted video to a specific distribution platform, means for automatically extracting important scenes from the video being viewed in real time and immediately notifying and displaying them on the user's display device, and means for generating explanations for the important scenes automatically extracted from the video being viewed in real time and automatically providing the explanations using a generation AI, thereby enabling users to immediately understand the significance of important scenes without missing them while viewing.
[0455] "Video data" refers to digital files of moving images that have been shot or created.
[0456] A "server" refers to a computer system that provides content and services in a network environment.
[0457] "Upload" refers to the operation of sending and saving data from a user's device to a server.
[0458] "Analysis" refers to the process of examining and evaluating uploaded video data to recognize specific elements or scenes.
[0459] An "important scene" refers to a scene or event in video data that should be of particular interest.
[0460] "Extraction" refers to the process of extracting specific scenes or information based on the analysis results.
[0461] "Description" refers to information that includes commentary and impressions about a particular scene.
[0462] "Overlay" refers to the process of displaying the generated explanation on a specific scene of the original video data.
[0463] "Delivery format" refers to the video specifications and formats supported by a particular platform.
[0464] "Transmission" refers to the act of delivering the converted video to a specific platform via a network.
[0465] "Real-time" refers to processing immediately at the time the video is being captured or streamed.
[0466] "Display device" refers to a hardware device that enables a user to visually receive information.
[0467] "Notification" refers to an operation that notifies a user of specific information or events.
[0468] "Generative AI" refers to artificial intelligence technology that automatically creates text and audio data based on visual and audio information.
[0469] The system of this invention uses generation AI to automatically generate important scenes and explanations from videos, creating high-quality clipped videos. In this system, the user, server, and device each play an important role.
[0470] Program Description:
[0471] Hardware and software used:
[0472] Hardware
[0473] Display devices (e.g., smart glasses, head-mounted displays)
[0474] Server (high performance computer system)
[0475] software
[0476] OpenCV (library for image analysis)
[0477] PyTorch (an execution environment for generative AI models)
[0478] Transformers library (models for natural language generation)
[0479] System behavior:
[0480] First, the user uploads the video data they have shot or created to the server via their device. The server then checks the meta information of the received video data and converts the video format as necessary. This process prepares the video data in a format that is easy for the generative AI model to analyze.
[0481] The server then analyzes the video data to detect specific words and scene changes. Based on this analysis, the system extracts important scenes from the video, such as a goal or the cheers of the crowd in a sports game.
[0482] For the extracted important scenes, a generative AI model is used to automatically generate explanations. The generated explanations are overlaid on the video data in a visually easy-to-understand format. This overlay process allows users to instantly understand the significance of important scenes.
[0483] The overlaid video of key moments is then converted into a specified distribution format. This distribution format is adjusted to meet the specifications required by specific platforms such as social media and video distribution sites. Finally, the converted video is sent to the specific distribution platform via a server.
[0484] Examples:
[0485] For example, consider a case where a user is watching a live stream of a soccer match with smart glasses. Any goal that the user missed will be automatically detected and extracted, and a description such as "Great goal!" will be automatically generated by the generative AI model. This description will be superimposed on the display of the smart glasses.
[0486] This allows users to instantly check important scenes they missed in real time and understand their significance.
[0487] Example prompt sentence:
[0488] "How can I automatically extract goal scenes from the live stream I'm watching and notify the user in a timely manner?"
[0489] In this way, the system of the present invention provides a powerful means for ensuring that users do not miss important scenes in a video in real time and for providing instant commentary on the scenes.
[0490] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0491] Step 1:
[0492] Uploading video data
[0493] A user uses a terminal to upload video data that has been shot or created to a server.
[0494] Input: Locally stored video file
[0495] Output: Video data stored on the server
[0496] Specific operation: The user selects video data through the application interface and uploads the data by pressing the send button. The server saves the received data in storage.
[0497] Step 2:
[0498] Check and convert meta information
[0499] The server checks the meta information (e.g., resolution, frame rate, format) of the received video data and performs format conversion if necessary.
[0500] Input: Video data stored on the server
[0501] Output: Video data converted for analysis
[0502] Specific operation: The server analyzes the video metadata and converts the resolution and format as necessary to make it suitable for analysis. For example, converting an AVI file to MP4 format.
[0503] Step 3:
[0504] Video data analysis
[0505] The server passes the video data prepared for analysis to the generative AI model and performs video analysis.
[0506] Input: Format converted video data
[0507] Output: Identification information of important scenes as analysis results
[0508] Specific operation: The server uses OpenCV and PyTorch to analyze each video frame, and the generative AI model recognizes specific scenes (e.g., goal scenes).
[0509] Step 4:
[0510] Extraction of important scenes
[0511] Based on the analysis results, the server extracts important scenes from the video.
[0512] Input: Analysis results from a generative AI model
[0513] Output: Extracted frames and timestamps of key scenes
[0514] Specific operation: The server identifies important frames such as "goal scenes" and "cheering from the crowd" from the analysis results of the generative AI model and extracts their timestamp information.
[0515] Step 5:
[0516] Generate Description
[0517] The server automatically generates explanations for the extracted important scenes using a generative AI model.
[0518] Input: Extracted important scenes
[0519] Output: Generated description text
[0520] What it does: The server uses the Transformers library to generate a description for each key moment, such as "Amazing goal!" or "Great play!"
[0521] Step 6:
[0522] Description overlay
[0523] The server overlays the generated explanation onto the video data.
[0524] Input: frames of key scenes and generated explanatory text
[0525] Output: Video of important scenes with overlaid explanations
[0526] Specific operation: The server uses OpenCV to overlay the generated explanatory text on specific frames and process it into a visually easy-to-understand form.
[0527] Step 7:
[0528] Conversion to delivery format
[0529] The server converts the superimposed important scene video into a specified distribution format.
[0530] Input: Video of key scenes with overlaid descriptions
[0531] Output: Video file converted to distribution format
[0532] Specific operation: The server adjusts the video resolution, aspect ratio, codec, etc. according to the platform specifications and converts it to the appropriate format.
[0533] Step 8:
[0534] Sending videos
[0535] The server sends the converted video file to a specific distribution platform.
[0536] Input: Video files converted to a distribution format
[0537] Output: Video published on a distribution platform
[0538] Specific operation: The server uploads the video through the specified API or endpoint and automatically publishes it.
[0539] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0540] The system of this invention uses a generative AI and an emotion engine to automatically generate highlight scenes and comments from videos and create high-quality clipped videos. In this system, the user, server, and emotion engine each play an important role.
[0541] Program processing
[0542] Uploading videos
[0543] Users upload video files to the server using their devices (PCs, smartphones, tablets, etc.). The uploaded video files are then stored in the server's storage.
[0544] Preparing for video analysis
[0545] The server checks the metadata of the uploaded video file and converts the format as needed, making it easier for the generation AI to analyze.
[0546] Video Analysis
[0547] The server passes the prepared video file to the AI generator and requests it to analyze the video. The AI analyzes the video in detail and identifies important scenes (highlight scenes) based on specific keywords and visual changes.
[0548] Highlight Scene Extraction and User Emotion Recognition
[0549] The generation AI extracts multiple highlight scenes based on the analysis results. At the same time, the emotion engine recognizes and collects the user's emotions while watching. The user's emotional data is reflected in the selection of highlight scenes.
[0550] Comment Generation
[0551] The server takes into account the user's emotional data obtained from the emotion engine and adds comments automatically generated by the AI to the highlight scenes. The comments are generated using natural language generation technology, including explanations and impressions of the scenes.
[0552] Generate cropped video
[0553] The server combines the extracted highlights with the generated comments to create a clipped video, during which the generated comments are added as text overlays to the video.
[0554] Format Conversion
[0555] The server then converts the resulting cropped video into the format of the specific platform (e.g., social networking site, video streaming site, etc.), which may include changing the aspect ratio and resolution of the video.
[0556] delivery
[0557] The server then distributes the optimized clipped video to a specific platform, allowing users to enjoy high-quality clipped videos on that platform. The distributed video also reflects the user's emotional information, making it more likely to resonate with viewers.
[0558] Specific examples
[0559] For example, consider the case where a user uploads a video of a soccer match from their smartphone to a server. When the user uploads the video file of the match, the server receives the file and converts the format if necessary. The server then passes the video file to a generation AI, which detects the score of the match, the cheers of the crowd, and so on.
[0560] At the same time, the emotion engine recognizes the user's emotions in real time as they watch and collects them as data. For example, if a user becomes excited at a goal, that emotional data is reflected in the generation AI. The generation AI then extracts particularly noteworthy scenes from the game. The server then generates comments such as "Amazing goal!" or "Great play!" Comments and highlight selection that reflect the emotional data evoke even stronger empathy in viewers.
[0561] The server combines these highlights and comments to create a clipped video that users can easily watch.The server then converts the video into a vertical format suitable for social media platforms and automatically posts it.As a result, users can quickly obtain high-quality clipped videos that reflect emotional information and share them on social media.
[0562] In this way, the system of the present invention utilizes advanced analysis of video content as well as user emotional data to automatically generate and efficiently distribute more attractive and relatable clipped videos.
[0563] The processing flow will be explained below.
[0564] Step 1:
[0565] The user uploads a video file to the server using the device. The user opens a dedicated upload screen, selects the video file, and clicks the send button.
[0566] Step 2:
[0567] The server receives the video file and stores it in temporary storage. The server checks the metadata of the video file, recording the file size and format. The server saves the received video file in the specified directory.
[0568] Step 3:
[0569] The server checks the format of the video file and converts it if necessary. If the file format is not a standard format (e.g. MP4), a conversion tool is used to convert it to the appropriate format.
[0570] Step 4:
[0571] The server calls an API to pass the prepared video file to the generation AI and starts the analysis job. The server records the analysis progress as a log and monitors the progress while the generation AI performs the analysis.
[0572] Step 5:
[0573] The generative AI analyzes the content of the video and extracts important scenes (highlight scenes) based on specific keywords (e.g., scores and cheers in sports), scene changes, and visual emphasis.
[0574] Step 6:
[0575] The emotion engine recognizes the user's emotions in real time while watching a video and collects them as data. User emotion data is acquired from facial expressions, voice, input information, etc. while watching a video.
[0576] Step 7:
[0577] The server again requests the AI to automatically generate comments for the highlight scenes obtained from the AI, taking into account the user's emotional data obtained from the emotion engine.
[0578] Step 8:
[0579] The server combines the highlight scenes and the generated comments to create a clipped video. Using a video editing library, the server clips out the specified highlight scenes and adds the generated comments to the video as a text overlay.
[0580] Step 9:
[0581] The server converts the resulting cropped video to the format of the specific platform, converting the video aspect ratio and resolution to a platform-appropriate format (e.g., vertical, 1080x1920).
[0582] Step 10:
[0583] The server distributes the converted video to a specific platform, and posts the clipped video to a specified account using the platform's API, such as a social networking site. This allows users to quickly obtain high-quality clipped videos that reflect their emotions.
[0584] Example 2
[0585] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0586] Conventional video editing systems require users to manually extract highlights and add comments, which requires time and effort. It is also difficult to achieve sophisticated editing that reflects the user's emotions, making it difficult to elicit empathy from viewers. Furthermore, format conversion and distribution for different platforms must be done manually, resulting in inefficiency.
[0587] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for uploading video files, means for analyzing video files, means for extracting highlight scenes based on the analysis results, means for generating comments for the extracted highlight scenes, means for recognizing and collecting user emotions, means for generating comments taking into account the collected emotion data, means for overlaying the generated comments on a video, means for converting the overlaid highlight video into a format for a specific platform, and means for delivering the converted video to a specific platform. This makes it possible to automatically generate high-quality clipped videos that reflect user emotions and efficiently deliver them to multiple platforms.
[0588] "Means for uploading video files" refers to the function that allows users to use their devices to send and save video files to a server.
[0589] The "means for analyzing video files" is a function that allows the server to automatically analyze the contents of uploaded video files.
[0590] The "means for extracting highlight scenes based on the analysis results" is a function for extracting specific important scenes based on the content analysis of the video.
[0591] The "means for generating a comment for the extracted highlight scene" is a function for automatically generating a comment in accordance with the extracted highlight scene.
[0592] "Means for recognizing and collecting user emotions" is a function that detects the user's emotions while watching in real time and collects them as data.
[0593] The "means for generating comments taking into consideration collected emotional data" is a function for generating comments that evoke greater empathy based on the user's emotional data.
[0594] The "means for overlaying generated comments on video" is a function for displaying generated comments overlaid on highlight scenes.
[0595] The "means for converting the overlaid highlight video into a format for a specific platform" is a function for converting the edited video into a format suitable for the standards of each platform.
[0596] "Means for distributing the converted video to a specific platform" refers to the function of uploading and distributing the format-converted video to platforms such as social media and video distribution sites.
[0597] The following describes an embodiment of the present invention. The system of the present invention utilizes a generation AI and an emotion engine to automatically generate highlight scenes and comments from videos and create high-quality clipped videos.
[0598] Uploading a video file
[0599] Users use their devices (PCs, smartphones, tablets, etc.) to upload video files to the server. They access the system's web application, log in, select the video file, and click the upload button. This video file is then saved in the server's storage. Specifically, a cloud storage service (e.g., Amazon S3) may be used.
[0600] Preparing for video analysis
[0601] The server receives the uploaded video file and checks its metadata. If necessary, it converts the video format to make it easier for the AI to analyze. For example, it uses software called FFmpeg to convert the original video format to H.264.
[0602] Video Analysis
[0603] The server passes the prepared video file to the generation AI, which uses, for example, the OpenAI API. The server inputs a prompt to the generation AI, such as "I want you to detect soccer goal scenes." The generation AI then performs a detailed analysis of specific keywords and visual changes in the video, outputting the timestamp and characteristics of each scene.
[0604] Highlight Scene Extraction and User Emotion Recognition
[0605] The generation AI extracts multiple highlight scenes based on the analysis results. At the same time, the emotion engine recognizes and collects the user's emotions in real time as they watch. This emotion recognition is performed using tools such as EmoPy. As the user watches the video, the server reads the user's facial expressions through the device's webcam. For example, if the user smiles during a goal-scoring scene, that emotional data is recorded.
[0606] Comment Generation
[0607] The server passes the user's emotional data obtained from the emotion engine to the generation AI. The generation AI generates comments such as "Incredible goal!" or "Great play!" based on data such as "This scene made the user smile." Specifically, natural language generation technology is often used. For example, the GPT-4 model is used.
[0608] Generate cropped video
[0609] The server combines the extracted highlights with the generated comments to create a clipped video, which is then added as a text overlay to the video. The video is edited using software such as Adobe Premiere Pro.
[0610] Format Conversion
[0611] The server then converts the resulting cut video into the format of the specific platform (e.g., social media, video streaming site, etc.), using tools such as FFmpeg to change the aspect ratio and resolution of the video.
[0612] delivery
[0613] The server automatically posts the optimized clip to a specific platform, often via the YouTube API or a social networking service API. Users receive a notification, watch the video, and share it on social media.
[0614] In this way, it is possible to automatically generate high-quality clipped videos that reflect the user's emotions and efficiently distribute them across multiple platforms, significantly reducing the time-consuming and labor-intensive work that was previously done manually, and providing video content that evokes empathy among viewers.
[0615] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0616] Step 1:
[0617] A user accesses the system's web application using a device (PC, smartphone, tablet, etc.) and logs in. Next, the user selects a video file and clicks the upload button. The input is the video file selected by the user, and the output is the video data sent to the server. The server stores this file in cloud storage (e.g., Amazon S3). The specific operation is an HTTP POST request.
[0618] Step 2:
[0619] The server receives the uploaded video file and checks the video file's metadata (resolution, format, length, etc.). If necessary, it uses FFmpeg to convert the original video format to H.264 format. The input is the original uploaded video file, and the output is the format-converted video file. Specific operations involve the execution of FFmpeg commands.
[0620] Step 3:
[0621] The server passes the converted video file to the generation AI (e.g., OpenAI's API). The input is the format-converted video file and a prompt statement such as, "I want you to detect soccer goal scenes." The output is the analysis results (timestamp and features of the goal scene) returned by the generation AI. Specific operations include sending an API request and a response.
[0622] Step 4:
[0623] The generation AI extracts multiple highlight scenes based on the analysis results. The server receives these and uses an emotion engine (e.g., EmoPy) to recognize and collect the user's emotions in real time while watching. The input is the analysis results and the user's real-time video (facial expression data), and the output is a highlight scene list including the user's emotional data. Specific operation uses a data stream via WebSocket.
[0624] Step 5:
[0625] The server passes the emotion data obtained from the emotion engine to the generation AI. Based on data such as "This scene made the user smile," the generation AI automatically generates comments such as "Incredible goal!" or "Great play!" The input is the user's emotion data and a list of highlight scenes, and the output is the generated comments. Specific operations include sending and responding to API requests.
[0626] Step 6:
[0627] The server combines the extracted highlight scenes and the generated comments to create a cut-out video. This process uses Adobe Premiere Pro scripts or similar video editing software. The input is the highlight scene list and the generated comments, and the output is the completed cut-out video. Specific operations include timeline editing and command execution.
[0628] Step 7:
[0629] The server uses FFmpeg to convert the completed cut video into the format of a specific platform (e.g., social networking site, video distribution site). The input is the edited cut video, and the output is a format-converted video file. Specifically, the FFmpeg command is executed again.
[0630] Step 8:
[0631] The server automatically posts clipped videos optimized for specific platforms via API. The input is the format-converted video file and the platform information for posting, and the output is the URL of the uploaded video. Specifically, a request is made to the API endpoint of each platform. The user receives this URL notification, watches it, and shares it with many people.
[0632] (Application example 2)
[0633] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0634] Conventional video analysis systems simply extract highlight scenes based on specific keywords or visual changes, without considering the user's emotions while watching. This resulted in the generation of videos that failed to appeal to the user's emotions and did not evoke empathy. Furthermore, the generation of comments was also uniform, which was an issue that did not fully reflect the user's emotions.
[0635] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0636] In this invention, the server includes means for uploading video files, means for analyzing video files, means for extracting highlight scenes based on the analysis results, means for generating comments for the extracted highlight scenes, means for overlaying the generated comments on the video, means for converting the overlaid highlight video into a format for a specific platform, means for distributing the converted video to a specific content distribution service, means for sensing a user's emotions while watching in real time and collecting emotion data, and means for generating comments based on the emotion data. This makes it possible to generate highlight videos and comments that reflect the user's emotions while watching and that evoke greater empathy.
[0637] The "means for uploading video files" is a function that allows a user to send video files from their own device to the server and store them in the server's storage.
[0638] "Means for analyzing video files" refers to technology that examines the video in detail and extracts necessary information in order to understand the content of the uploaded video file.
[0639] "Means for extracting highlight scenes" refers to a technology that identifies important or noteworthy scenes as a result of video analysis and extracts them.
[0640] The "means for generating comments" is a function that uses a generative AI model to create natural language sentences for extracted highlight scenes.
[0641] "Video overlay means" refers to a technique for displaying generated comments or other information overlaid on video frames.
[0642] The "means for converting into a format" refers to a technology for changing the resolution, aspect ratio, etc. of the generated highlight video in order to optimally display it on a specific platform.
[0643] "Means of distribution" refers to the function of automatically sending and publishing the converted video to a specified content distribution service or social networking site.
[0644] "Means for sensing emotions in real time" refers to technology that detects the emotional state of a user in real time from facial expressions, voice, etc. while the user is watching a video.
[0645] The "means for collecting emotion data" refers to a technology for recording detected emotion information of a user as data and using it for subsequent analysis and processing.
[0646] "Means for generating comments based on emotional data" refers to a technology in which a generative AI model automatically creates comments that are appropriate to the user's emotions, taking into account collected emotional data.
[0647] The system for implementing the present invention includes a terminal, a server, a generative AI model, and an emotion engine. The specific operation of this system will be described in detail below.
[0648] The terminal is a device for users to upload video files, and can be a smartphone, tablet, or PC. Using this terminal, users upload video files they have taken. The video files are sent to the server and stored in the server's storage.
[0649] The server analyzes the uploaded video files and extracts important scenes (highlight scenes) based on the analysis results. This analysis uses a generative AI model to analyze the video content in detail. The server also converts the format of the video file as needed to make it easier to analyze.
[0650] Next, the server uses an emotion engine to sense the user's emotions in real time while watching the video and collects emotion data. This emotion data is acquired while the user is watching the video, and the user's emotional state is detected from their facial expressions, voice, gestures, etc. The data obtained by the emotion engine is sent to a generative AI model and reflected in the selection of highlight scenes and comment generation.
[0651] The generative AI model takes into account the collected emotional data and automatically generates appropriate comments for the extracted highlight scenes. These comments are in line with the video content and the user's emotions, aiming to evoke empathy in the viewer. The generated comments are added as a text overlay to the video.
[0652] The overlaid highlight footage is then converted into a format suitable for the specific platform (e.g. YouTube, Instagram, etc.) The server performs this conversion, adjusting the video resolution, aspect ratio, etc.
[0653] Finally, the server automatically distributes the converted highlight video to the specified content distribution service or social networking site, allowing users to easily publish and share the video.
[0654] As a concrete example, consider the case where a user uploads a video of a soccer match from their smartphone to a server. When the user uploads the video file of the match, the server receives the file and converts the format as necessary. The server then passes the video file to a generative AI model, which detects things like goal scenes and crowd cheers. At the same time, as the user watches, the emotion engine recognizes the user's emotions in real time and collects them as data. For example, if the user becomes excited at a goal, that emotional data is reflected in the generative AI model, which selects the scene as particularly noteworthy. The server then generates comments such as "Amazing goal!" or "Great play!"
[0655] Below are some example prompts to input to a generative AI model:
[0656] "Please extract the goal scenes and excitement of the crowd from this soccer match video, generate optimal comments based on the user's emotional data, and create a cut-out video."
[0657] In this way, the system can enhance the user's viewing experience and generate high-quality videos that evoke stronger empathy in the viewer.
[0658] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0659] Step 1:
[0660] The server receives video files uploaded by users from their devices and stores them in the server's storage. The input is the video file from the user's device, and the output is the video file stored in the server storage. In this step, file reception and storage operations are performed.
[0661] Step 2:
[0662] The server checks the metadata of the stored video file and performs format conversion if necessary. The input is the stored video file, and the output is a video file converted into a format that is easy for the generative AI model to analyze. In this step, metadata is checked and format conversion is performed.
[0663] Step 3:
[0664] The server passes the converted video file to the generative AI model, which then performs a detailed analysis of the video content. The input is the converted video file, and the output is the highlight scenes identified as a result of the video analysis. In this step, detailed video analysis is performed using the generative AI model.
[0665] Step 4:
[0666] The server uses an emotion engine to sense the user's emotions in real time while the user is watching and collects emotion data. The input is the user's facial expressions and voice while watching, and the output is the emotion data collected in real time. In this step, emotion recognition and data collection take place.
[0667] Step 5:
[0668] The server uses a generative AI model based on the collected emotion data to generate comments for the highlight scenes. The input is emotion data and the highlight scenes, and the output is the generated comments. In this step, emotion data is analyzed and comments are generated.
[0669] Step 6:
[0670] The server adds the generated comments to the video as a text overlay. The input is the highlight scene and the generated comments, and the output is the highlight video with the comments overlayed. In this step, the text overlay is added.
[0671] Step 7:
[0672] The server converts the overlaid highlight video into a specific platform format. The input is the highlight video with overlaid comments, and the output is a video format optimized for the specific platform. In this step, the format conversion takes place.
[0673] Step 8:
[0674] The server automatically distributes the converted highlight video to the specified content distribution service or social networking site. The input is a video file converted for a specific platform, and the output is a distributed and published highlight video. In this step, the video is distributed and published.
[0675] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0676] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0677] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0678] [Third embodiment]
[0679] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0680] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0681] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0682] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0683] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0684] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0685] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0686] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0687] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0688] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0689] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0690] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0691] The system of this invention uses AI to automatically generate highlight scenes and comments from videos, creating high-quality clipped videos. In this system, the user, server, and device each play an important role.
[0692] Program processing
[0693] Uploading videos
[0694] Users upload video files to the server using their devices (PCs, smartphones, tablets, etc.). The uploaded video files are then stored in the server's storage.
[0695] Preparing for video analysis
[0696] The server checks the metadata of the uploaded video file and converts the format as needed, making it easier for the generation AI to analyze.
[0697] Video Analysis
[0698] The server passes the prepared video file to the AI generator and requests it to analyze the video. The AI analyzes the video in detail and identifies important scenes (highlight scenes) based on specific keywords and visual changes.
[0699] Highlight Scene Extraction
[0700] Based on the analysis results, the generation AI extracts multiple highlight scenes, which constitute the parts that are particularly interesting to the user.
[0701] Comment Generation
[0702] The server automatically adds comments to the highlight scenes using AI, which are generated using natural language generation technology and include explanations and impressions of the scenes.
[0703] Generate cropped video
[0704] The server combines the extracted highlights with the generated comments to create a clipped video, during which the generated comments are added as text overlays to the video.
[0705] Format Conversion
[0706] The server then converts the resulting cropped video into the format of the specific platform (e.g., social networking site, video streaming site, etc.), which may include changing the aspect ratio and resolution of the video.
[0707] delivery
[0708] The server delivers the optimized cut-out video to a specific platform, allowing users to enjoy high-quality cut-out video on that platform.
[0709] Specific examples
[0710] For example, consider the case where a user uploads a video of a soccer match from their smartphone to a server. When the user uploads the video file of the match, the server receives the file and converts the format if necessary. The server then passes the video file to a generation AI, which detects the score of the match, the cheers of the crowd, and so on.
[0711] The generation AI extracts particularly noteworthy scenes from a match and generates comments for them, such as "Amazing goal!" or "Great play!" The server combines these highlight scenes and comments to create a clipped video that users can easily watch. The server then converts the video into a vertical format suitable for social media platforms and automatically posts it. As a result, users can quickly obtain an appealing clipped video and share it on social media.
[0712] In this way, the system of the present invention provides a powerful means for users to easily create and distribute high-quality clipped videos through advanced analysis and automatic processing of video content.
[0713] The processing flow will be explained below.
[0714] Step 1:
[0715] The user uploads a video file to the server using the device. The user opens a dedicated upload screen, selects the video file, and clicks the send button.
[0716] Step 2:
[0717] The server receives the video file and stores it in temporary storage. The server checks the metadata of the video file, recording the file size and format. The server saves the received video file in the specified directory.
[0718] Step 3:
[0719] The server checks the format of the video file and converts it if necessary. If the file format is not a standard format (e.g. MP4), a conversion tool is used to convert it to the appropriate format.
[0720] Step 4:
[0721] The server calls an API to pass the prepared video file to the generation AI and starts the analysis job. The server records the analysis progress as a log and monitors the progress while the generation AI performs the analysis.
[0722] Step 5:
[0723] The generative AI analyzes the content of the video and extracts important scenes (highlight scenes) based on specific keywords (e.g., scores and cheers in sports), scene changes, and visual emphasis.
[0724] Step 6:
[0725] The server then requests the AI to automatically generate comments for the highlight scenes obtained from the AI. The AI then uses natural language generation technology to generate comments appropriate for the highlight scenes.
[0726] Step 7:
[0727] The server combines the highlight scenes and the generated comments to create a clipped video. Using a video editing library, the server clips out the specified highlight scenes and adds the generated comments to the video as a text overlay.
[0728] Step 8:
[0729] The server converts the resulting cropped video to the format of the specific platform, converting the video aspect ratio and resolution to a platform-appropriate format (e.g., vertical, 1080x1920).
[0730] Step 9:
[0731] The server distributes the converted video to a specific platform, and posts the clipped video to a specified account using the platform's API, such as a social networking site. This allows users to obtain high-quality clipped videos in a short amount of time.
[0732] Example 1
[0733] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0734] Analyzing and editing video data takes time and effort, especially extracting highlight scenes, generating captions for those scenes, and converting them into formats suitable for specific platforms. This makes it difficult for users to quickly create and distribute high-quality clipped videos.
[0735] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0736] In this invention, the server includes means for uploading video data, means for checking metadata of the video data and converting the format as necessary, means for analyzing the video data, means for extracting important scenes based on the analysis results, means for generating explanatory text for the extracted important scenes, means for overlaying the generated explanatory text on the video, means for converting the overlaid important video into a format for a specific platform, and means for delivering the converted video to the specific platform. This allows users to easily upload video data and quickly generate and deliver high-quality cut-out videos.
[0737] "Video data" means digital files containing visual and audio information.
[0738] "Metadata" is additional information that accompanies video data, including technical details such as file format, resolution, frame rate, bit rate, etc.
[0739] "Format conversion" is the process of converting a file from one digital format to another digital format.
[0740] "Analysis" refers to the detailed investigation and analysis of digital information, and when applied to video data in particular, refers to the evaluation of the video and audio content using specialized algorithms.
[0741] An "important scene" is a particularly noteworthy portion of video data, and refers to a scene with prominent visual or audio characteristics or a scene that is likely to be of great interest to the user.
[0742] A "description" is a sentence generated for a specific scene, which expresses the content and meaning of the scene in words.
[0743] "Text overlay" is a technique for displaying text information overlaid on specific scenes in a video.
[0744] "Platform" refers to an online system or environment for distributing video data and enabling users to view it.
[0745] "Distribution" refers to the operation or process of delivering content to users, especially over the Internet.
[0746] This invention relates to a system that uses generative AI to automatically generate highlight scenes and comments from videos and create high-quality clips. By using this system, users can easily upload videos, create clips based on the analysis results, and distribute them to specific platforms.
[0747] First, a user uploads video data to the server using a device such as a PC, smartphone, or tablet. For example, a user selects a soccer game video taken with a smartphone and presses the "upload" button. The video data is then transferred to the server. A progress bar is displayed to indicate the progress.
[0748] The server then checks the metadata of the uploaded video data and converts the format if necessary. For example, video data in .avi format is converted to .mp4 format, which is easier for the generative AI model to analyze.
[0749] Once the server has completed preparing the video data, it passes the data to the generative AI and requests it to analyze it. The generative AI model is then sent a prompt such as, "Please identify the highlight scenes in this video." At this time, the generative AI model detects specific keywords (e.g., "goal" or "cheers") and visual changes in the video to identify important scenes.
[0750] The AI generator creates a list of important scenes (highlights) based on the analysis results. For example, you can set the time range and scene content, such as "00:05:20-00:05:40: The goal scene" or "00:10:10-00:10:30: The crowd cheering scene."
[0751] The server then asks the AI to generate a description for the highlight scene. For example, it sends a prompt such as, "Please generate a description for this scene." The AI generates a comment for the "goal scene" such as, "A great shot to score!"
[0752] The server then combines the extracted highlights with the generated captions to create a short video clip. For example, it adds a text overlay to a goal highlight, saying "Great shot, score!", and saves it as a short video clip.
[0753] The server then converts the finished cropped video to the format for the specific platform, for example converting it from a 16:9 landscape format for YouTube to a 9:16 portrait format for Instagram, and adjusting the resolution to 1080p.
[0754] Finally, the server distributes the optimized video to a specific platform, such as Instagram, where it automatically posts the video to the user's account, automatically setting the video title and tags.
[0755] In this way, the user, server, and device work together to easily upload video data and quickly generate and distribute high-quality clipped videos. A specific example of a prompt is as follows:
[0756] "Identify the highlights in this video and generate comments for the parts that include goals and crowd cheers."
[0757] "Extract important scenes from a soccer game video taken with a smartphone and generate a text overlay for them."
[0758] "Analyze the most exciting moments in the game, add engaging commentary and create clips."
[0759] This allows users to easily upload video data and quickly generate and distribute high-quality clipped videos.
[0760] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0761] Step 1:
[0762] Users upload video data to a server using devices such as PCs, smartphones, and tablets. For example, when a user selects a video they have taken with their smartphone and presses the "upload" button, the video data is transferred to the server. The input is the video data uploaded from the user's device, and the output is the video data saved on the server.
[0763] Step 2:
[0764] The server checks the metadata of the uploaded video data and determines whether format conversion is necessary. For example, the server analyzes the video file format, resolution, frame rate, etc., and converts it to .mp4 format if necessary. The input is the original video data stored on the server, and the output is video data converted into a format that is easy for the generative AI to analyze.
[0765] Step 3:
[0766] The server passes the prepared video data to the generation AI and requests analysis using a prompt. For example, the prompt "Please identify the highlight scene in this video" is sent to the generation AI. The generation AI analyzes the video data and detects specific keywords and visual changes. The input is the video data stored on the server and the prompt, and the output is a list of candidate highlight scenes as the analysis result.
[0767] Step 4:
[0768] The generation AI lists important scenes based on the analysis results. For example, you can set the time range and scene content, such as "00:05:20-00:05:40 goal scene" or "00:10:10-00:10:30 crowd cheering scene." The input is the analysis result data, and the output is a list containing detailed time ranges and scene content.
[0769] Step 5:
[0770] The server requests the generation AI to generate a description for each highlight scene. For example, it sends a prompt saying, "Please generate a description for this scene." The generation AI generates an appropriate description for that scene. The input is the highlight scene data and the prompt, and the output is a description corresponding to each scene.
[0771] Step 6:
[0772] The server combines the extracted highlight scenes with the generated explanatory text as a text overlay to create a clipped video. For example, a text overlay saying "Great shot and score!" is added to a goal scene. The input is the video data of the highlight scenes and explanatory text, and the output is an edited clipped video.
[0773] Step 7:
[0774] The server then converts the completed cropped video to a format suitable for the specific platform, e.g., converting from 16:9 to 9:16 portrait format for YouTube or Instagram. The input is the cropped video, and the output is the converted video for the specific platform.
[0775] Step 8:
[0776] The server delivers the optimized video to a specific platform, for example, using the Instagram API to automatically post the video to a user's account. The input is the platform-optimized video data, and the output is the delivered video.
[0777] In this way, through each processing step of the system, users can easily generate high-quality cutout videos and share them on specific platforms.
[0778] (Application example 1)
[0779] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0780] Conventional video distribution systems lack the technology to automatically identify and extract important scenes that users may have missed while watching, and notify the viewing device of the missed scenes at the appropriate time. This causes users to miss important scenes, resulting in a poor viewing experience. Furthermore, there is a lack of a means to immediately provide appropriate explanations for extracted important scenes, making it difficult for viewers to understand the significance of the scenes.
[0781] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0782] In this invention, the server includes means for uploading video data, means for analyzing the video data, means for extracting important scenes based on the analysis results, means for generating explanations for the extracted important scenes, means for overlaying the generated explanations on the video, means for converting the overlaid important scene video into a specific distribution format, means for transmitting the converted video to a specific distribution platform, means for automatically extracting important scenes from the video being viewed in real time and immediately notifying and displaying them on the user's display device, and means for generating explanations for the important scenes automatically extracted from the video being viewed in real time and automatically providing the explanations using a generation AI, thereby enabling users to immediately understand the significance of important scenes without missing them while viewing.
[0783] "Video data" refers to digital files of moving images that have been shot or created.
[0784] A "server" refers to a computer system that provides content and services in a network environment.
[0785] "Upload" refers to the operation of sending and saving data from a user's device to a server.
[0786] "Analysis" refers to the process of examining and evaluating uploaded video data to recognize specific elements or scenes.
[0787] An "important scene" refers to a scene or event in video data that should be of particular interest.
[0788] "Extraction" refers to the process of extracting specific scenes or information based on the analysis results.
[0789] "Description" refers to information that includes commentary and impressions about a particular scene.
[0790] "Overlay" refers to the process of displaying the generated explanation on a specific scene of the original video data.
[0791] "Delivery format" refers to the video specifications and formats supported by a particular platform.
[0792] "Transmission" refers to the act of delivering the converted video to a specific platform via a network.
[0793] "Real-time" refers to processing immediately at the time the video is being captured or streamed.
[0794] "Display device" refers to a hardware device that enables a user to visually receive information.
[0795] "Notification" refers to an operation that notifies a user of specific information or events.
[0796] "Generative AI" refers to artificial intelligence technology that automatically creates text and audio data based on visual and audio information.
[0797] The system of this invention uses generation AI to automatically generate important scenes and explanations from videos, creating high-quality clipped videos. In this system, the user, server, and device each play an important role.
[0798] Program Description:
[0799] Hardware and software used:
[0800] Hardware
[0801] Display devices (e.g., smart glasses, head-mounted displays)
[0802] Server (high performance computer system)
[0803] software
[0804] OpenCV (library for image analysis)
[0805] PyTorch (an execution environment for generative AI models)
[0806] Transformers library (models for natural language generation)
[0807] System behavior:
[0808] First, the user uploads the video data they have shot or created to the server via their device. The server then checks the meta information of the received video data and converts the video format as necessary. This process prepares the video data in a format that is easy for the generative AI model to analyze.
[0809] The server then analyzes the video data to detect specific words and scene changes. Based on this analysis, the system extracts important scenes from the video, such as a goal or the cheers of the crowd in a sports game.
[0810] For the extracted important scenes, a generative AI model is used to automatically generate explanations. The generated explanations are overlaid on the video data in a visually easy-to-understand format. This overlay process allows users to instantly understand the significance of important scenes.
[0811] The overlaid video of key moments is then converted into a specified distribution format. This distribution format is adjusted to meet the specifications required by specific platforms such as social media and video distribution sites. Finally, the converted video is sent to the specific distribution platform via a server.
[0812] Examples:
[0813] For example, consider a case where a user is watching a live stream of a soccer match with smart glasses. Any goal that the user missed will be automatically detected and extracted, and a description such as "Great goal!" will be automatically generated by the generative AI model. This description will be superimposed on the display of the smart glasses.
[0814] This allows users to instantly check important scenes they missed in real time and understand their significance.
[0815] Example prompt sentence:
[0816] "How can I automatically extract goal scenes from the live stream I'm watching and notify the user in a timely manner?"
[0817] In this way, the system of the present invention provides a powerful means for ensuring that users do not miss important scenes in a video in real time and for providing instant commentary on the scenes.
[0818] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0819] Step 1:
[0820] Uploading video data
[0821] A user uses a terminal to upload video data that has been shot or created to a server.
[0822] Input: Locally stored video file
[0823] Output: Video data stored on the server
[0824] Specific operation: The user selects video data through the application interface and uploads the data by pressing the send button. The server saves the received data in storage.
[0825] Step 2:
[0826] Check and convert meta information
[0827] The server checks the meta information (e.g., resolution, frame rate, format) of the received video data and performs format conversion if necessary.
[0828] Input: Video data stored on the server
[0829] Output: Video data converted for analysis
[0830] Specific operation: The server analyzes the video metadata and converts the resolution and format as necessary to make it suitable for analysis. For example, converting an AVI file to MP4 format.
[0831] Step 3:
[0832] Video data analysis
[0833] The server passes the video data prepared for analysis to the generative AI model and performs video analysis.
[0834] Input: Format converted video data
[0835] Output: Identification information of important scenes as analysis results
[0836] Specific operation: The server uses OpenCV and PyTorch to analyze each video frame, and the generative AI model recognizes specific scenes (e.g., goal scenes).
[0837] Step 4:
[0838] Extraction of important scenes
[0839] Based on the analysis results, the server extracts important scenes from the video.
[0840] Input: Analysis results from a generative AI model
[0841] Output: Extracted frames and timestamps of key scenes
[0842] Specific operation: The server identifies important frames such as "goal scenes" and "cheering from the crowd" from the analysis results of the generative AI model and extracts their timestamp information.
[0843] Step 5:
[0844] Generate Description
[0845] The server automatically generates explanations for the extracted important scenes using a generative AI model.
[0846] Input: Extracted important scenes
[0847] Output: Generated description text
[0848] What it does: The server uses the Transformers library to generate a description for each key moment, such as "Amazing goal!" or "Great play!"
[0849] Step 6:
[0850] Description overlay
[0851] The server overlays the generated explanation onto the video data.
[0852] Input: frames of key scenes and generated explanatory text
[0853] Output: Video of important scenes with overlaid explanations
[0854] Specific operation: The server uses OpenCV to overlay the generated explanatory text on specific frames and process it into a visually easy-to-understand form.
[0855] Step 7:
[0856] Conversion to delivery format
[0857] The server converts the superimposed important scene video into a specified distribution format.
[0858] Input: Video of key scenes with overlaid descriptions
[0859] Output: Video file converted to distribution format
[0860] Specific operation: The server adjusts the video resolution, aspect ratio, codec, etc. according to the platform specifications and converts it to the appropriate format.
[0861] Step 8:
[0862] Sending videos
[0863] The server sends the converted video file to a specific distribution platform.
[0864] Input: Video files converted to a distribution format
[0865] Output: Video published on a distribution platform
[0866] Specific operation: The server uploads the video through the specified API or endpoint and automatically publishes it.
[0867] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0868] The system of this invention uses a generative AI and an emotion engine to automatically generate highlight scenes and comments from videos and create high-quality clipped videos. In this system, the user, server, and emotion engine each play an important role.
[0869] Program processing
[0870] Uploading videos
[0871] Users upload video files to the server using their devices (PCs, smartphones, tablets, etc.). The uploaded video files are then stored in the server's storage.
[0872] Preparing for video analysis
[0873] The server checks the metadata of the uploaded video file and converts the format as needed, making it easier for the generation AI to analyze.
[0874] Video Analysis
[0875] The server passes the prepared video file to the AI generator and requests it to analyze the video. The AI analyzes the video in detail and identifies important scenes (highlight scenes) based on specific keywords and visual changes.
[0876] Highlight Scene Extraction and User Emotion Recognition
[0877] The generation AI extracts multiple highlight scenes based on the analysis results. At the same time, the emotion engine recognizes and collects the user's emotions while watching. The user's emotional data is reflected in the selection of highlight scenes.
[0878] Comment Generation
[0879] The server takes into account the user's emotional data obtained from the emotion engine and adds comments automatically generated by the AI to the highlight scenes. The comments are generated using natural language generation technology, including explanations and impressions of the scenes.
[0880] Generate cropped video
[0881] The server combines the extracted highlights with the generated comments to create a clipped video, during which the generated comments are added as text overlays to the video.
[0882] Format Conversion
[0883] The server then converts the resulting cropped video into the format of the specific platform (e.g., social networking site, video streaming site, etc.), which may include changing the aspect ratio and resolution of the video.
[0884] delivery
[0885] The server then distributes the optimized clipped video to a specific platform, allowing users to enjoy high-quality clipped videos on that platform. The distributed video also reflects the user's emotional information, making it more likely to resonate with viewers.
[0886] Specific examples
[0887] For example, consider the case where a user uploads a video of a soccer match from their smartphone to a server. When the user uploads the video file of the match, the server receives the file and converts the format if necessary. The server then passes the video file to a generation AI, which detects the score of the match, the cheers of the crowd, and so on.
[0888] At the same time, the emotion engine recognizes the user's emotions in real time as they watch and collects them as data. For example, if a user becomes excited at a goal, that emotional data is reflected in the generation AI. The generation AI then extracts particularly noteworthy scenes from the game. The server then generates comments such as "Amazing goal!" or "Great play!" Comments and highlight selection that reflect the emotional data evoke even stronger empathy in viewers.
[0889] The server combines these highlights and comments to create a clipped video that users can easily watch.The server then converts the video into a vertical format suitable for social media platforms and automatically posts it.As a result, users can quickly obtain high-quality clipped videos that reflect emotional information and share them on social media.
[0890] In this way, the system of the present invention utilizes advanced analysis of video content as well as user emotional data to automatically generate and efficiently distribute more attractive and relatable clipped videos.
[0891] The processing flow will be explained below.
[0892] Step 1:
[0893] The user uploads a video file to the server using the device. The user opens a dedicated upload screen, selects the video file, and clicks the send button.
[0894] Step 2:
[0895] The server receives the video file and stores it in temporary storage. The server checks the metadata of the video file, recording the file size and format. The server saves the received video file in the specified directory.
[0896] Step 3:
[0897] The server checks the format of the video file and converts it if necessary. If the file format is not a standard format (e.g. MP4), a conversion tool is used to convert it to the appropriate format.
[0898] Step 4:
[0899] The server calls an API to pass the prepared video file to the generation AI and starts the analysis job. The server records the analysis progress as a log and monitors the progress while the generation AI performs the analysis.
[0900] Step 5:
[0901] The generative AI analyzes the content of the video and extracts important scenes (highlight scenes) based on specific keywords (e.g., scores and cheers in sports), scene changes, and visual emphasis.
[0902] Step 6:
[0903] The emotion engine recognizes the user's emotions in real time while watching a video and collects them as data. User emotion data is acquired from facial expressions, voice, input information, etc. while watching a video.
[0904] Step 7:
[0905] The server again requests the AI to automatically generate comments for the highlight scenes obtained from the AI, taking into account the user's emotional data obtained from the emotion engine.
[0906] Step 8:
[0907] The server combines the highlight scenes and the generated comments to create a clipped video. Using a video editing library, the server clips out the specified highlight scenes and adds the generated comments to the video as a text overlay.
[0908] Step 9:
[0909] The server converts the resulting cropped video to the format of the specific platform, converting the video aspect ratio and resolution to a platform-appropriate format (e.g., vertical, 1080x1920).
[0910] Step 10:
[0911] The server distributes the converted video to a specific platform, and posts the clipped video to a specified account using the platform's API, such as a social networking site. This allows users to quickly obtain high-quality clipped videos that reflect their emotions.
[0912] Example 2
[0913] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0914] Conventional video editing systems require users to manually extract highlights and add comments, which requires time and effort. It is also difficult to achieve sophisticated editing that reflects the user's emotions, making it difficult to elicit empathy from viewers. Furthermore, format conversion and distribution for different platforms must be done manually, resulting in inefficiency.
[0915] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for uploading video files, means for analyzing video files, means for extracting highlight scenes based on the analysis results, means for generating comments for the extracted highlight scenes, means for recognizing and collecting user emotions, means for generating comments taking into account the collected emotion data, means for overlaying the generated comments on a video, means for converting the overlaid highlight video into a format for a specific platform, and means for delivering the converted video to a specific platform. This makes it possible to automatically generate high-quality clipped videos that reflect user emotions and efficiently deliver them to multiple platforms.
[0916] "Means for uploading video files" refers to the function that allows users to use their devices to send and save video files to a server.
[0917] The "means for analyzing video files" is a function that allows the server to automatically analyze the contents of uploaded video files.
[0918] The "means for extracting highlight scenes based on the analysis results" is a function for extracting specific important scenes based on the content analysis of the video.
[0919] The "means for generating a comment for the extracted highlight scene" is a function for automatically generating a comment in accordance with the extracted highlight scene.
[0920] "Means for recognizing and collecting user emotions" is a function that detects the user's emotions while watching in real time and collects them as data.
[0921] The "means for generating comments taking into consideration collected emotional data" is a function for generating comments that evoke greater empathy based on the user's emotional data.
[0922] The "means for overlaying generated comments on video" is a function for displaying generated comments overlaid on highlight scenes.
[0923] The "means for converting the overlaid highlight video into a format for a specific platform" is a function for converting the edited video into a format suitable for the standards of each platform.
[0924] "Means for distributing the converted video to a specific platform" refers to the function of uploading and distributing the format-converted video to platforms such as social media and video distribution sites.
[0925] The following describes an embodiment of the present invention. The system of the present invention utilizes a generation AI and an emotion engine to automatically generate highlight scenes and comments from videos and create high-quality clipped videos.
[0926] Uploading a video file
[0927] Users use their devices (PCs, smartphones, tablets, etc.) to upload video files to the server. They access the system's web application, log in, select the video file, and click the upload button. This video file is then saved in the server's storage. Specifically, a cloud storage service (e.g., Amazon S3) may be used.
[0928] Preparing for video analysis
[0929] The server receives the uploaded video file and checks its metadata. If necessary, it converts the video format to make it easier for the AI to analyze. For example, it uses software called FFmpeg to convert the original video format to H.264.
[0930] Video Analysis
[0931] The server passes the prepared video file to the generation AI, which uses, for example, the OpenAI API. The server inputs a prompt to the generation AI, such as "I want you to detect soccer goal scenes." The generation AI then performs a detailed analysis of specific keywords and visual changes in the video, outputting the timestamp and characteristics of each scene.
[0932] Highlight Scene Extraction and User Emotion Recognition
[0933] The generation AI extracts multiple highlight scenes based on the analysis results. At the same time, the emotion engine recognizes and collects the user's emotions in real time as they watch. This emotion recognition is performed using tools such as EmoPy. As the user watches the video, the server reads the user's facial expressions through the device's webcam. For example, if the user smiles during a goal-scoring scene, that emotional data is recorded.
[0934] Comment Generation
[0935] The server passes the user's emotional data obtained from the emotion engine to the generation AI. The generation AI generates comments such as "Incredible goal!" or "Great play!" based on data such as "This scene made the user smile." Specifically, natural language generation technology is often used. For example, the GPT-4 model is used.
[0936] Generate cropped video
[0937] The server combines the extracted highlights with the generated comments to create a clipped video, which is then added as a text overlay to the video. The video is edited using software such as Adobe Premiere Pro.
[0938] Format Conversion
[0939] The server then converts the resulting cut video into the format of the specific platform (e.g., social media, video streaming site, etc.), using tools such as FFmpeg to change the aspect ratio and resolution of the video.
[0940] delivery
[0941] The server automatically posts the optimized clip to a specific platform, often via the YouTube API or a social networking service API. Users receive a notification, watch the video, and share it on social media.
[0942] In this way, it is possible to automatically generate high-quality clipped videos that reflect the user's emotions and efficiently distribute them across multiple platforms, significantly reducing the time-consuming and labor-intensive work that was previously done manually, and providing video content that evokes empathy among viewers.
[0943] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0944] Step 1:
[0945] A user accesses the system's web application using a device (PC, smartphone, tablet, etc.) and logs in. Next, the user selects a video file and clicks the upload button. The input is the video file selected by the user, and the output is the video data sent to the server. The server stores this file in cloud storage (e.g., Amazon S3). The specific operation is an HTTP POST request.
[0946] Step 2:
[0947] The server receives the uploaded video file and checks the video file's metadata (resolution, format, length, etc.). If necessary, it uses FFmpeg to convert the original video format to H.264 format. The input is the original uploaded video file, and the output is the format-converted video file. Specific operations involve the execution of FFmpeg commands.
[0948] Step 3:
[0949] The server passes the converted video file to the generation AI (e.g., OpenAI's API). The input is the format-converted video file and a prompt statement such as, "I want you to detect soccer goal scenes." The output is the analysis results (timestamp and features of the goal scene) returned by the generation AI. Specific operations include sending an API request and a response.
[0950] Step 4:
[0951] The generation AI extracts multiple highlight scenes based on the analysis results. The server receives these and uses an emotion engine (e.g., EmoPy) to recognize and collect the user's emotions in real time while watching. The input is the analysis results and the user's real-time video (facial expression data), and the output is a highlight scene list including the user's emotional data. Specific operation uses a data stream via WebSocket.
[0952] Step 5:
[0953] The server passes the emotion data obtained from the emotion engine to the generation AI. Based on data such as "This scene made the user smile," the generation AI automatically generates comments such as "Incredible goal!" or "Great play!" The input is the user's emotion data and a list of highlight scenes, and the output is the generated comments. Specific operations include sending and responding to API requests.
[0954] Step 6:
[0955] The server combines the extracted highlight scenes and the generated comments to create a cut-out video. This process uses Adobe Premiere Pro scripts or similar video editing software. The input is the highlight scene list and the generated comments, and the output is the completed cut-out video. Specific operations include timeline editing and command execution.
[0956] Step 7:
[0957] The server uses FFmpeg to convert the completed cut video into the format of a specific platform (e.g., social networking site, video distribution site). The input is the edited cut video, and the output is a format-converted video file. Specifically, the FFmpeg command is executed again.
[0958] Step 8:
[0959] The server automatically posts clipped videos optimized for specific platforms via API. The input is the format-converted video file and the platform information for posting, and the output is the URL of the uploaded video. Specifically, a request is made to the API endpoint of each platform. The user receives this URL notification, watches it, and shares it with many people.
[0960] (Application example 2)
[0961] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0962] Conventional video analysis systems simply extract highlight scenes based on specific keywords or visual changes, without considering the user's emotions while watching. This resulted in the generation of videos that failed to appeal to the user's emotions and did not evoke empathy. Furthermore, the generation of comments was also uniform, which was an issue that did not fully reflect the user's emotions.
[0963] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0964] In this invention, the server includes means for uploading video files, means for analyzing video files, means for extracting highlight scenes based on the analysis results, means for generating comments for the extracted highlight scenes, means for overlaying the generated comments on the video, means for converting the overlaid highlight video into a format for a specific platform, means for distributing the converted video to a specific content distribution service, means for sensing a user's emotions while watching in real time and collecting emotion data, and means for generating comments based on the emotion data. This makes it possible to generate highlight videos and comments that reflect the user's emotions while watching and that evoke greater empathy.
[0965] The "means for uploading video files" is a function that allows a user to send video files from their own device to the server and store them in the server's storage.
[0966] "Means for analyzing video files" refers to technology that examines the video in detail and extracts necessary information in order to understand the content of the uploaded video file.
[0967] "Means for extracting highlight scenes" refers to a technology that identifies important or noteworthy scenes as a result of video analysis and extracts them.
[0968] The "means for generating comments" is a function that uses a generative AI model to create natural language sentences for extracted highlight scenes.
[0969] "Video overlay means" refers to a technique for displaying generated comments or other information overlaid on video frames.
[0970] The "means for converting into a format" refers to a technology for changing the resolution, aspect ratio, etc. of the generated highlight video in order to optimally display it on a specific platform.
[0971] "Means of distribution" refers to the function of automatically sending and publishing the converted video to a specified content distribution service or social networking site.
[0972] "Means for sensing emotions in real time" refers to technology that detects the emotional state of a user in real time from facial expressions, voice, etc. while the user is watching a video.
[0973] The "means for collecting emotion data" refers to a technology for recording detected emotion information of a user as data and using it for subsequent analysis and processing.
[0974] "Means for generating comments based on emotional data" refers to a technology in which a generative AI model automatically creates comments that are appropriate to the user's emotions, taking into account collected emotional data.
[0975] The system for implementing the present invention includes a terminal, a server, a generative AI model, and an emotion engine. The specific operation of this system will be described in detail below.
[0976] The terminal is a device for users to upload video files, and can be a smartphone, tablet, or PC. Using this terminal, users upload video files they have taken. The video files are sent to the server and stored in the server's storage.
[0977] The server analyzes the uploaded video files and extracts important scenes (highlight scenes) based on the analysis results. This analysis uses a generative AI model to analyze the video content in detail. The server also converts the format of the video file as needed to make it easier to analyze.
[0978] Next, the server uses an emotion engine to sense the user's emotions in real time while watching the video and collects emotion data. This emotion data is acquired while the user is watching the video, and the user's emotional state is detected from their facial expressions, voice, gestures, etc. The data obtained by the emotion engine is sent to a generative AI model and reflected in the selection of highlight scenes and comment generation.
[0979] The generative AI model takes into account the collected emotional data and automatically generates appropriate comments for the extracted highlight scenes. These comments are in line with the video content and the user's emotions, aiming to evoke empathy in the viewer. The generated comments are added as a text overlay to the video.
[0980] The overlaid highlight footage is then converted into a format suitable for the specific platform (e.g. YouTube, Instagram, etc.) The server performs this conversion, adjusting the video resolution, aspect ratio, etc.
[0981] Finally, the server automatically distributes the converted highlight video to the specified content distribution service or social networking site, allowing users to easily publish and share the video.
[0982] As a concrete example, consider the case where a user uploads a video of a soccer match from their smartphone to a server. When the user uploads the video file of the match, the server receives the file and converts the format as necessary. The server then passes the video file to a generative AI model, which detects things like goal scenes and crowd cheers. At the same time, as the user watches, the emotion engine recognizes the user's emotions in real time and collects them as data. For example, if the user becomes excited at a goal, that emotional data is reflected in the generative AI model, which selects the scene as particularly noteworthy. The server then generates comments such as "Amazing goal!" or "Great play!"
[0983] Below are some example prompts to input to a generative AI model:
[0984] "Please extract the goal scenes and excitement of the crowd from this soccer match video, generate optimal comments based on the user's emotional data, and create a cut-out video."
[0985] In this way, the system can enhance the user's viewing experience and generate high-quality videos that evoke stronger empathy in the viewer.
[0986] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0987] Step 1:
[0988] The server receives video files uploaded by users from their devices and stores them in the server's storage. The input is the video file from the user's device, and the output is the video file stored in the server storage. In this step, file reception and storage operations are performed.
[0989] Step 2:
[0990] The server checks the metadata of the stored video file and performs format conversion if necessary. The input is the stored video file, and the output is a video file converted into a format that is easy for the generative AI model to analyze. In this step, metadata is checked and format conversion is performed.
[0991] Step 3:
[0992] The server passes the converted video file to the generative AI model, which then performs a detailed analysis of the video content. The input is the converted video file, and the output is the highlight scenes identified as a result of the video analysis. In this step, detailed video analysis is performed using the generative AI model.
[0993] Step 4:
[0994] The server uses an emotion engine to sense the user's emotions in real time while the user is watching and collects emotion data. The input is the user's facial expressions and voice while watching, and the output is the emotion data collected in real time. In this step, emotion recognition and data collection take place.
[0995] Step 5:
[0996] The server uses a generative AI model based on the collected emotion data to generate comments for the highlight scenes. The input is emotion data and the highlight scenes, and the output is the generated comments. In this step, emotion data is analyzed and comments are generated.
[0997] Step 6:
[0998] The server adds the generated comments to the video as a text overlay. The input is the highlight scene and the generated comments, and the output is the highlight video with the comments overlayed. In this step, the text overlay is added.
[0999] Step 7:
[1000] The server converts the overlaid highlight video into a specific platform format. The input is the highlight video with overlaid comments, and the output is a video format optimized for the specific platform. In this step, the format conversion takes place.
[1001] Step 8:
[1002] The server automatically distributes the converted highlight video to the specified content distribution service or social networking site. The input is a video file converted for a specific platform, and the output is a distributed and published highlight video. In this step, the video is distributed and published.
[1003] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1004] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1005] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1006] [Fourth embodiment]
[1007] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1008] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1009] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1010] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1011] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1012] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1013] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1014] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1015] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1016] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1017] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1018] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1019] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1020] The system of this invention uses AI to automatically generate highlight scenes and comments from videos, creating high-quality clipped videos. In this system, the user, server, and device each play an important role.
[1021] Program processing
[1022] Uploading videos
[1023] Users upload video files to the server using their devices (PCs, smartphones, tablets, etc.). The uploaded video files are then stored in the server's storage.
[1024] Preparing for video analysis
[1025] The server checks the metadata of the uploaded video file and converts the format as needed, making it easier for the generation AI to analyze.
[1026] Video Analysis
[1027] The server passes the prepared video file to the AI generator and requests it to analyze the video. The AI analyzes the video in detail and identifies important scenes (highlight scenes) based on specific keywords and visual changes.
[1028] Highlight Scene Extraction
[1029] Based on the analysis results, the generation AI extracts multiple highlight scenes, which constitute the parts that are particularly interesting to the user.
[1030] Comment Generation
[1031] The server automatically adds comments to the highlight scenes using AI, which are generated using natural language generation technology and include explanations and impressions of the scenes.
[1032] Generate cropped video
[1033] The server combines the extracted highlights with the generated comments to create a clipped video, during which the generated comments are added as text overlays to the video.
[1034] Format Conversion
[1035] The server then converts the resulting cropped video into the format of the specific platform (e.g., social networking site, video streaming site, etc.), which may include changing the aspect ratio and resolution of the video.
[1036] delivery
[1037] The server delivers the optimized cut-out video to a specific platform, allowing users to enjoy high-quality cut-out video on that platform.
[1038] Specific examples
[1039] For example, consider the case where a user uploads a video of a soccer match from their smartphone to a server. When the user uploads the video file of the match, the server receives the file and converts the format if necessary. The server then passes the video file to a generation AI, which detects the score of the match, the cheers of the crowd, and so on.
[1040] The generation AI extracts particularly noteworthy scenes from a match and generates comments for them, such as "Amazing goal!" or "Great play!" The server combines these highlight scenes and comments to create a clipped video that users can easily watch. The server then converts the video into a vertical format suitable for social media platforms and automatically posts it. As a result, users can quickly obtain an appealing clipped video and share it on social media.
[1041] In this way, the system of the present invention provides a powerful means for users to easily create and distribute high-quality clipped videos through advanced analysis and automatic processing of video content.
[1042] The processing flow will be explained below.
[1043] Step 1:
[1044] The user uploads a video file to the server using the device. The user opens a dedicated upload screen, selects the video file, and clicks the send button.
[1045] Step 2:
[1046] The server receives the video file and stores it in temporary storage. The server checks the metadata of the video file, recording the file size and format. The server saves the received video file in the specified directory.
[1047] Step 3:
[1048] The server checks the format of the video file and converts it if necessary. If the file format is not a standard format (e.g. MP4), a conversion tool is used to convert it to the appropriate format.
[1049] Step 4:
[1050] The server calls an API to pass the prepared video file to the generation AI and starts the analysis job. The server records the analysis progress as a log and monitors the progress while the generation AI performs the analysis.
[1051] Step 5:
[1052] The generative AI analyzes the content of the video and extracts important scenes (highlight scenes) based on specific keywords (e.g., scores and cheers in sports), scene changes, and visual emphasis.
[1053] Step 6:
[1054] The server then requests the AI to automatically generate comments for the highlight scenes obtained from the AI. The AI then uses natural language generation technology to generate comments appropriate for the highlight scenes.
[1055] Step 7:
[1056] The server combines the highlight scenes and the generated comments to create a clipped video. Using a video editing library, the server clips out the specified highlight scenes and adds the generated comments to the video as a text overlay.
[1057] Step 8:
[1058] The server converts the resulting cropped video to the format of the specific platform, converting the video aspect ratio and resolution to a platform-appropriate format (e.g., vertical, 1080x1920).
[1059] Step 9:
[1060] The server distributes the converted video to a specific platform, and posts the clipped video to a specified account using the platform's API, such as a social networking site. This allows users to obtain high-quality clipped videos in a short amount of time.
[1061] Example 1
[1062] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1063] Analyzing and editing video data takes time and effort, especially extracting highlight scenes, generating captions for those scenes, and converting them into formats suitable for specific platforms. This makes it difficult for users to quickly create and distribute high-quality clipped videos.
[1064] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1065] In this invention, the server includes means for uploading video data, means for checking metadata of the video data and converting the format as necessary, means for analyzing the video data, means for extracting important scenes based on the analysis results, means for generating explanatory text for the extracted important scenes, means for overlaying the generated explanatory text on the video, means for converting the overlaid important video into a format for a specific platform, and means for delivering the converted video to the specific platform. This allows users to easily upload video data and quickly generate and deliver high-quality cut-out videos.
[1066] "Video data" means digital files containing visual and audio information.
[1067] "Metadata" is additional information that accompanies video data, including technical details such as file format, resolution, frame rate, bit rate, etc.
[1068] "Format conversion" is the process of converting a file from one digital format to another digital format.
[1069] "Analysis" refers to the detailed investigation and analysis of digital information, and when applied to video data in particular, refers to the evaluation of the video and audio content using specialized algorithms.
[1070] An "important scene" is a particularly noteworthy portion of video data, and refers to a scene with prominent visual or audio characteristics or a scene that is likely to be of great interest to the user.
[1071] A "description" is a sentence generated for a specific scene, which expresses the content and meaning of the scene in words.
[1072] "Text overlay" is a technique for displaying text information overlaid on specific scenes in a video.
[1073] "Platform" refers to an online system or environment for distributing video data and enabling users to view it.
[1074] "Distribution" refers to the operation or process of delivering content to users, especially over the Internet.
[1075] This invention relates to a system that uses generative AI to automatically generate highlight scenes and comments from videos and create high-quality clips. By using this system, users can easily upload videos, create clips based on the analysis results, and distribute them to specific platforms.
[1076] First, a user uploads video data to the server using a device such as a PC, smartphone, or tablet. For example, a user selects a soccer game video taken with a smartphone and presses the "upload" button. The video data is then transferred to the server. A progress bar is displayed to indicate the progress.
[1077] The server then checks the metadata of the uploaded video data and converts the format if necessary. For example, video data in .avi format is converted to .mp4 format, which is easier for the generative AI model to analyze.
[1078] Once the server has completed preparing the video data, it passes the data to the generative AI and requests it to analyze it. The generative AI model is then sent a prompt such as, "Please identify the highlight scenes in this video." At this time, the generative AI model detects specific keywords (e.g., "goal" or "cheers") and visual changes in the video to identify important scenes.
[1079] The AI generator creates a list of important scenes (highlights) based on the analysis results. For example, you can set the time range and scene content, such as "00:05:20-00:05:40: The goal scene" or "00:10:10-00:10:30: The crowd cheering scene."
[1080] The server then asks the AI to generate a description for the highlight scene. For example, it sends a prompt such as, "Please generate a description for this scene." The AI generates a comment for the "goal scene" such as, "A great shot to score!"
[1081] The server then combines the extracted highlights with the generated captions to create a short video clip. For example, it adds a text overlay to a goal highlight, saying "Great shot, score!", and saves it as a short video clip.
[1082] The server then converts the finished cropped video to the format for the specific platform, for example converting it from a 16:9 landscape format for YouTube to a 9:16 portrait format for Instagram, and adjusting the resolution to 1080p.
[1083] Finally, the server distributes the optimized video to a specific platform, such as Instagram, where it automatically posts the video to the user's account, automatically setting the video title and tags.
[1084] In this way, the user, server, and device work together to easily upload video data and quickly generate and distribute high-quality clipped videos. A specific example of a prompt is as follows:
[1085] "Identify the highlights in this video and generate comments for the parts that include goals and crowd cheers."
[1086] "Extract important scenes from a soccer game video taken with a smartphone and generate a text overlay for them."
[1087] "Analyze the most exciting moments in the game, add engaging commentary and create clips."
[1088] This allows users to easily upload video data and quickly generate and distribute high-quality clipped videos.
[1089] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1090] Step 1:
[1091] Users upload video data to a server using devices such as PCs, smartphones, and tablets. For example, when a user selects a video they have taken with their smartphone and presses the "upload" button, the video data is transferred to the server. The input is the video data uploaded from the user's device, and the output is the video data saved on the server.
[1092] Step 2:
[1093] The server checks the metadata of the uploaded video data and determines whether format conversion is necessary. For example, the server analyzes the video file format, resolution, frame rate, etc., and converts it to .mp4 format if necessary. The input is the original video data stored on the server, and the output is video data converted into a format that is easy for the generative AI to analyze.
[1094] Step 3:
[1095] The server passes the prepared video data to the generation AI and requests analysis using a prompt. For example, the prompt "Please identify the highlight scene in this video" is sent to the generation AI. The generation AI analyzes the video data and detects specific keywords and visual changes. The input is the video data stored on the server and the prompt, and the output is a list of candidate highlight scenes as the analysis result.
[1096] Step 4:
[1097] The generation AI lists important scenes based on the analysis results. For example, you can set the time range and scene content, such as "00:05:20-00:05:40 goal scene" or "00:10:10-00:10:30 crowd cheering scene." The input is the analysis result data, and the output is a list containing detailed time ranges and scene content.
[1098] Step 5:
[1099] The server requests the generation AI to generate a description for each highlight scene. For example, it sends a prompt saying, "Please generate a description for this scene." The generation AI generates an appropriate description for that scene. The input is the highlight scene data and the prompt, and the output is a description corresponding to each scene.
[1100] Step 6:
[1101] The server combines the extracted highlight scenes with the generated explanatory text as a text overlay to create a clipped video. For example, a text overlay saying "Great shot and score!" is added to a goal scene. The input is the video data of the highlight scenes and explanatory text, and the output is an edited clipped video.
[1102] Step 7:
[1103] The server then converts the completed cropped video to a format suitable for the specific platform, e.g., converting from 16:9 to 9:16 portrait format for YouTube or Instagram. The input is the cropped video, and the output is the converted video for the specific platform.
[1104] Step 8:
[1105] The server delivers the optimized video to a specific platform, for example, using the Instagram API to automatically post the video to a user's account. The input is the platform-optimized video data, and the output is the delivered video.
[1106] In this way, through each processing step of the system, users can easily generate high-quality cutout videos and share them on specific platforms.
[1107] (Application example 1)
[1108] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1109] Conventional video distribution systems lack the technology to automatically identify and extract important scenes that users may have missed while watching, and notify the viewing device of the missed scenes at the appropriate time. This causes users to miss important scenes, resulting in a poor viewing experience. Furthermore, there is a lack of a means to immediately provide appropriate explanations for extracted important scenes, making it difficult for viewers to understand the significance of the scenes.
[1110] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1111] In this invention, the server includes means for uploading video data, means for analyzing the video data, means for extracting important scenes based on the analysis results, means for generating explanations for the extracted important scenes, means for overlaying the generated explanations on the video, means for converting the overlaid important scene video into a specific distribution format, means for transmitting the converted video to a specific distribution platform, means for automatically extracting important scenes from the video being viewed in real time and immediately notifying and displaying them on the user's display device, and means for generating explanations for the important scenes automatically extracted from the video being viewed in real time and automatically providing the explanations using a generation AI, thereby enabling users to immediately understand the significance of important scenes without missing them while viewing.
[1112] "Video data" refers to digital files of moving images that have been shot or created.
[1113] A "server" refers to a computer system that provides content and services in a network environment.
[1114] "Upload" refers to the operation of sending and saving data from a user's device to a server.
[1115] "Analysis" refers to the process of examining and evaluating uploaded video data to recognize specific elements or scenes.
[1116] An "important scene" refers to a scene or event in video data that should be of particular interest.
[1117] "Extraction" refers to the process of extracting specific scenes or information based on the analysis results.
[1118] "Description" refers to information that includes commentary and impressions about a particular scene.
[1119] "Overlay" refers to the process of displaying the generated explanation on a specific scene of the original video data.
[1120] "Delivery format" refers to the video specifications and formats supported by a particular platform.
[1121] "Transmission" refers to the act of delivering the converted video to a specific platform via a network.
[1122] "Real-time" refers to processing immediately at the time the video is being captured or streamed.
[1123] "Display device" refers to a hardware device that enables a user to visually receive information.
[1124] "Notification" refers to an operation that notifies a user of specific information or events.
[1125] "Generative AI" refers to artificial intelligence technology that automatically creates text and audio data based on visual and audio information.
[1126] The system of this invention uses generation AI to automatically generate important scenes and explanations from videos, creating high-quality clipped videos. In this system, the user, server, and device each play an important role.
[1127] Program Description:
[1128] Hardware and software used:
[1129] Hardware
[1130] Display devices (e.g., smart glasses, head-mounted displays)
[1131] Server (high performance computer system)
[1132] software
[1133] OpenCV (library for image analysis)
[1134] PyTorch (an execution environment for generative AI models)
[1135] Transformers library (models for natural language generation)
[1136] System behavior:
[1137] First, the user uploads the video data they have shot or created to the server via their device. The server then checks the meta information of the received video data and converts the video format as necessary. This process prepares the video data in a format that is easy for the generative AI model to analyze.
[1138] The server then analyzes the video data to detect specific words and scene changes. Based on this analysis, the system extracts important scenes from the video, such as a goal or the cheers of the crowd in a sports game.
[1139] For the extracted important scenes, a generative AI model is used to automatically generate explanations. The generated explanations are overlaid on the video data in a visually easy-to-understand format. This overlay process allows users to instantly understand the significance of important scenes.
[1140] The overlaid video of key moments is then converted into a specified distribution format. This distribution format is adjusted to meet the specifications required by specific platforms such as social media and video distribution sites. Finally, the converted video is sent to the specific distribution platform via a server.
[1141] Examples:
[1142] For example, consider a case where a user is watching a live stream of a soccer match with smart glasses. Any goal that the user missed will be automatically detected and extracted, and a description such as "Great goal!" will be automatically generated by the generative AI model. This description will be superimposed on the display of the smart glasses.
[1143] This allows users to instantly check important scenes they missed in real time and understand their significance.
[1144] Example prompt sentence:
[1145] "How can I automatically extract goal scenes from the live stream I'm watching and notify the user in a timely manner?"
[1146] In this way, the system of the present invention provides a powerful means for ensuring that users do not miss important scenes in a video in real time and for providing instant commentary on the scenes.
[1147] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1148] Step 1:
[1149] Uploading video data
[1150] A user uses a terminal to upload video data that has been shot or created to a server.
[1151] Input: Locally stored video file
[1152] Output: Video data stored on the server
[1153] Specific operation: The user selects video data through the application interface and uploads the data by pressing the send button. The server saves the received data in storage.
[1154] Step 2:
[1155] Check and convert meta information
[1156] The server checks the meta information (e.g., resolution, frame rate, format) of the received video data and performs format conversion if necessary.
[1157] Input: Video data stored on the server
[1158] Output: Video data converted for analysis
[1159] Specific operation: The server analyzes the video metadata and converts the resolution and format as necessary to make it suitable for analysis. For example, converting an AVI file to MP4 format.
[1160] Step 3:
[1161] Video data analysis
[1162] The server passes the video data prepared for analysis to the generative AI model and performs video analysis.
[1163] Input: Format converted video data
[1164] Output: Identification information of important scenes as analysis results
[1165] Specific operation: The server uses OpenCV and PyTorch to analyze each video frame, and the generative AI model recognizes specific scenes (e.g., goal scenes).
[1166] Step 4:
[1167] Extraction of important scenes
[1168] Based on the analysis results, the server extracts important scenes from the video.
[1169] Input: Analysis results from a generative AI model
[1170] Output: Extracted frames and timestamps of key scenes
[1171] Specific operation: The server identifies important frames such as "goal scenes" and "cheering from the crowd" from the analysis results of the generative AI model and extracts their timestamp information.
[1172] Step 5:
[1173] Generate Description
[1174] The server automatically generates explanations for the extracted important scenes using a generative AI model.
[1175] Input: Extracted important scenes
[1176] Output: Generated description text
[1177] What it does: The server uses the Transformers library to generate a description for each key moment, such as "Amazing goal!" or "Great play!"
[1178] Step 6:
[1179] Description overlay
[1180] The server overlays the generated explanation onto the video data.
[1181] Input: frames of key scenes and generated explanatory text
[1182] Output: Video of important scenes with overlaid explanations
[1183] Specific operation: The server uses OpenCV to overlay the generated explanatory text on specific frames and process it into a visually easy-to-understand form.
[1184] Step 7:
[1185] Conversion to delivery format
[1186] The server converts the superimposed important scene video into a specified distribution format.
[1187] Input: Video of key scenes with overlaid descriptions
[1188] Output: Video file converted to distribution format
[1189] Specific operation: The server adjusts the video resolution, aspect ratio, codec, etc. according to the platform specifications and converts it to the appropriate format.
[1190] Step 8:
[1191] Sending videos
[1192] The server sends the converted video file to a specific distribution platform.
[1193] Input: Video files converted to a distribution format
[1194] Output: Video published on a distribution platform
[1195] Specific operation: The server uploads the video through the specified API or endpoint and automatically publishes it.
[1196] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1197] The system of this invention uses a generative AI and an emotion engine to automatically generate highlight scenes and comments from videos and create high-quality clipped videos. In this system, the user, server, and emotion engine each play an important role.
[1198] Program processing
[1199] Uploading videos
[1200] Users upload video files to the server using their devices (PCs, smartphones, tablets, etc.). The uploaded video files are then stored in the server's storage.
[1201] Preparing for video analysis
[1202] The server checks the metadata of the uploaded video file and converts the format as needed, making it easier for the generation AI to analyze.
[1203] Video Analysis
[1204] The server passes the prepared video file to the AI generator and requests it to analyze the video. The AI analyzes the video in detail and identifies important scenes (highlight scenes) based on specific keywords and visual changes.
[1205] Highlight Scene Extraction and User Emotion Recognition
[1206] The generation AI extracts multiple highlight scenes based on the analysis results. At the same time, the emotion engine recognizes and collects the user's emotions while watching. The user's emotional data is reflected in the selection of highlight scenes.
[1207] Comment Generation
[1208] The server takes into account the user's emotional data obtained from the emotion engine and adds comments automatically generated by the AI to the highlight scenes. The comments are generated using natural language generation technology, including explanations and impressions of the scenes.
[1209] Generate cropped video
[1210] The server combines the extracted highlights with the generated comments to create a clipped video, during which the generated comments are added as text overlays to the video.
[1211] Format Conversion
[1212] The server then converts the resulting cropped video into the format of the specific platform (e.g., social networking site, video streaming site, etc.), which may include changing the aspect ratio and resolution of the video.
[1213] delivery
[1214] The server then distributes the optimized clipped video to a specific platform, allowing users to enjoy high-quality clipped videos on that platform. The distributed video also reflects the user's emotional information, making it more likely to resonate with viewers.
[1215] Specific examples
[1216] For example, consider the case where a user uploads a video of a soccer match from their smartphone to a server. When the user uploads the video file of the match, the server receives the file and converts the format if necessary. The server then passes the video file to a generation AI, which detects the score of the match, the cheers of the crowd, and so on.
[1217] At the same time, the emotion engine recognizes the user's emotions in real time as they watch and collects them as data. For example, if a user becomes excited at a goal, that emotional data is reflected in the generation AI. The generation AI then extracts particularly noteworthy scenes from the game. The server then generates comments such as "Amazing goal!" or "Great play!" Comments and highlight selection that reflect the emotional data evoke even stronger empathy in viewers.
[1218] The server combines these highlights and comments to create a clipped video that users can easily watch.The server then converts the video into a vertical format suitable for social media platforms and automatically posts it.As a result, users can quickly obtain high-quality clipped videos that reflect emotional information and share them on social media.
[1219] In this way, the system of the present invention utilizes advanced analysis of video content as well as user emotional data to automatically generate and efficiently distribute more attractive and relatable clipped videos.
[1220] The processing flow will be explained below.
[1221] Step 1:
[1222] The user uploads a video file to the server using the device. The user opens a dedicated upload screen, selects the video file, and clicks the send button.
[1223] Step 2:
[1224] The server receives the video file and stores it in temporary storage. The server checks the metadata of the video file, recording the file size and format. The server saves the received video file in the specified directory.
[1225] Step 3:
[1226] The server checks the format of the video file and converts it if necessary. If the file format is not a standard format (e.g. MP4), a conversion tool is used to convert it to the appropriate format.
[1227] Step 4:
[1228] The server calls an API to pass the prepared video file to the generation AI and starts the analysis job. The server records the analysis progress as a log and monitors the progress while the generation AI performs the analysis.
[1229] Step 5:
[1230] The generative AI analyzes the content of the video and extracts important scenes (highlight scenes) based on specific keywords (e.g., scores and cheers in sports), scene changes, and visual emphasis.
[1231] Step 6:
[1232] The emotion engine recognizes the user's emotions in real time while watching a video and collects them as data. User emotion data is acquired from facial expressions, voice, input information, etc. while watching a video.
[1233] Step 7:
[1234] The server again requests the AI to automatically generate comments for the highlight scenes obtained from the AI, taking into account the user's emotional data obtained from the emotion engine.
[1235] Step 8:
[1236] The server combines the highlight scenes and the generated comments to create a clipped video. Using a video editing library, the server clips out the specified highlight scenes and adds the generated comments to the video as a text overlay.
[1237] Step 9:
[1238] The server converts the resulting cropped video to the format of the specific platform, converting the video aspect ratio and resolution to a platform-appropriate format (e.g., vertical, 1080x1920).
[1239] Step 10:
[1240] The server distributes the converted video to a specific platform, and posts the clipped video to a specified account using the platform's API, such as a social networking site. This allows users to quickly obtain high-quality clipped videos that reflect their emotions.
[1241] Example 2
[1242] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1243] Conventional video editing systems require users to manually extract highlights and add comments, which requires time and effort. It is also difficult to achieve sophisticated editing that reflects the user's emotions, making it difficult to elicit empathy from viewers. Furthermore, format conversion and distribution for different platforms must be done manually, resulting in inefficiency.
[1244] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for uploading video files, means for analyzing video files, means for extracting highlight scenes based on the analysis results, means for generating comments for the extracted highlight scenes, means for recognizing and collecting user emotions, means for generating comments taking into account the collected emotion data, means for overlaying the generated comments on a video, means for converting the overlaid highlight video into a format for a specific platform, and means for delivering the converted video to a specific platform. This makes it possible to automatically generate high-quality clipped videos that reflect user emotions and efficiently deliver them to multiple platforms.
[1245] "Means for uploading video files" refers to the function that allows users to use their devices to send and save video files to a server.
[1246] The "means for analyzing video files" is a function that allows the server to automatically analyze the contents of uploaded video files.
[1247] The "means for extracting highlight scenes based on the analysis results" is a function for extracting specific important scenes based on the content analysis of the video.
[1248] The "means for generating a comment for the extracted highlight scene" is a function for automatically generating a comment in accordance with the extracted highlight scene.
[1249] "Means for recognizing and collecting user emotions" is a function that detects the user's emotions while watching in real time and collects them as data.
[1250] The "means for generating comments taking into consideration collected emotional data" is a function for generating comments that evoke greater empathy based on the user's emotional data.
[1251] The "means for overlaying generated comments on video" is a function for displaying generated comments overlaid on highlight scenes.
[1252] The "means for converting the overlaid highlight video into a format for a specific platform" is a function for converting the edited video into a format suitable for the standards of each platform.
[1253] "Means for distributing the converted video to a specific platform" refers to the function of uploading and distributing the format-converted video to platforms such as social media and video distribution sites.
[1254] The following describes an embodiment of the present invention. The system of the present invention utilizes a generation AI and an emotion engine to automatically generate highlight scenes and comments from videos and create high-quality clipped videos.
[1255] Uploading a video file
[1256] Users use their devices (PCs, smartphones, tablets, etc.) to upload video files to the server. They access the system's web application, log in, select the video file, and click the upload button. This video file is then saved in the server's storage. Specifically, a cloud storage service (e.g., Amazon S3) may be used.
[1257] Preparing for video analysis
[1258] The server receives the uploaded video file and checks its metadata. If necessary, it converts the video format to make it easier for the AI to analyze. For example, it uses software called FFmpeg to convert the original video format to H.264.
[1259] Video Analysis
[1260] The server passes the prepared video file to the generation AI, which uses, for example, the OpenAI API. The server inputs a prompt to the generation AI, such as "I want you to detect soccer goal scenes." The generation AI then performs a detailed analysis of specific keywords and visual changes in the video, outputting the timestamp and characteristics of each scene.
[1261] Highlight Scene Extraction and User Emotion Recognition
[1262] The generation AI extracts multiple highlight scenes based on the analysis results. At the same time, the emotion engine recognizes and collects the user's emotions in real time as they watch. This emotion recognition is performed using tools such as EmoPy. As the user watches the video, the server reads the user's facial expressions through the device's webcam. For example, if the user smiles during a goal-scoring scene, that emotional data is recorded.
[1263] Comment Generation
[1264] The server passes the user's emotional data obtained from the emotion engine to the generation AI. The generation AI generates comments such as "Incredible goal!" or "Great play!" based on data such as "This scene made the user smile." Specifically, natural language generation technology is often used. For example, the GPT-4 model is used.
[1265] Generate cropped video
[1266] The server combines the extracted highlights with the generated comments to create a clipped video, which is then added as a text overlay to the video. The video is edited using software such as Adobe Premiere Pro.
[1267] Format Conversion
[1268] The server then converts the resulting cut video into the format of the specific platform (e.g., social media, video streaming site, etc.), using tools such as FFmpeg to change the aspect ratio and resolution of the video.
[1269] delivery
[1270] The server automatically posts the optimized clip to a specific platform, often via the YouTube API or a social networking service API. Users receive a notification, watch the video, and share it on social media.
[1271] In this way, it is possible to automatically generate high-quality clipped videos that reflect the user's emotions and efficiently distribute them across multiple platforms, significantly reducing the time-consuming and labor-intensive work that was previously done manually, and providing video content that evokes empathy among viewers.
[1272] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1273] Step 1:
[1274] A user accesses the system's web application using a device (PC, smartphone, tablet, etc.) and logs in. Next, the user selects a video file and clicks the upload button. The input is the video file selected by the user, and the output is the video data sent to the server. The server stores this file in cloud storage (e.g., Amazon S3). The specific operation is an HTTP POST request.
[1275] Step 2:
[1276] The server receives the uploaded video file and checks the video file's metadata (resolution, format, length, etc.). If necessary, it uses FFmpeg to convert the original video format to H.264 format. The input is the original uploaded video file, and the output is the format-converted video file. Specific operations involve the execution of FFmpeg commands.
[1277] Step 3:
[1278] The server passes the converted video file to the generation AI (e.g., OpenAI's API). The input is the format-converted video file and a prompt statement such as, "I want you to detect soccer goal scenes." The output is the analysis results (timestamp and features of the goal scene) returned by the generation AI. Specific operations include sending an API request and a response.
[1279] Step 4:
[1280] The generation AI extracts multiple highlight scenes based on the analysis results. The server receives these and uses an emotion engine (e.g., EmoPy) to recognize and collect the user's emotions in real time while watching. The input is the analysis results and the user's real-time video (facial expression data), and the output is a highlight scene list including the user's emotional data. Specific operation uses a data stream via WebSocket.
[1281] Step 5:
[1282] The server passes the emotion data obtained from the emotion engine to the generation AI. Based on data such as "This scene made the user smile," the generation AI automatically generates comments such as "Incredible goal!" or "Great play!" The input is the user's emotion data and a list of highlight scenes, and the output is the generated comments. Specific operations include sending and responding to API requests.
[1283] Step 6:
[1284] The server combines the extracted highlight scenes and the generated comments to create a cut-out video. This process uses Adobe Premiere Pro scripts or similar video editing software. The input is the highlight scene list and the generated comments, and the output is the completed cut-out video. Specific operations include timeline editing and command execution.
[1285] Step 7:
[1286] The server uses FFmpeg to convert the completed cut video into the format of a specific platform (e.g., social networking site, video distribution site). The input is the edited cut video, and the output is a format-converted video file. Specifically, the FFmpeg command is executed again.
[1287] Step 8:
[1288] The server automatically posts clipped videos optimized for specific platforms via API. The input is the format-converted video file and the platform information for posting, and the output is the URL of the uploaded video. Specifically, a request is made to the API endpoint of each platform. The user receives this URL notification, watches it, and shares it with many people.
[1289] (Application example 2)
[1290] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1291] Conventional video analysis systems simply extract highlight scenes based on specific keywords or visual changes, without considering the user's emotions while watching. This resulted in the generation of videos that failed to appeal to the user's emotions and did not evoke empathy. Furthermore, the generation of comments was also uniform, which was an issue that did not fully reflect the user's emotions.
[1292] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1293] In this invention, the server includes means for uploading video files, means for analyzing video files, means for extracting highlight scenes based on the analysis results, means for generating comments for the extracted highlight scenes, means for overlaying the generated comments on the video, means for converting the overlaid highlight video into a format for a specific platform, means for distributing the converted video to a specific content distribution service, means for sensing a user's emotions while watching in real time and collecting emotion data, and means for generating comments based on the emotion data. This makes it possible to generate highlight videos and comments that reflect the user's emotions while watching and that evoke greater empathy.
[1294] The "means for uploading video files" is a function that allows a user to send video files from their own device to the server and store them in the server's storage.
[1295] "Means for analyzing video files" refers to technology that examines the video in detail and extracts necessary information in order to understand the content of the uploaded video file.
[1296] "Means for extracting highlight scenes" refers to a technology that identifies important or noteworthy scenes as a result of video analysis and extracts them.
[1297] The "means for generating comments" is a function that uses a generative AI model to create natural language sentences for extracted highlight scenes.
[1298] "Video overlay means" refers to a technique for displaying generated comments or other information overlaid on video frames.
[1299] The "means for converting into a format" refers to a technology for changing the resolution, aspect ratio, etc. of the generated highlight video in order to optimally display it on a specific platform.
[1300] "Means of distribution" refers to the function of automatically sending and publishing the converted video to a specified content distribution service or social networking site.
[1301] "Means for sensing emotions in real time" refers to technology that detects the emotional state of a user in real time from facial expressions, voice, etc. while the user is watching a video.
[1302] The "means for collecting emotion data" refers to a technology for recording detected emotion information of a user as data and using it for subsequent analysis and processing.
[1303] "Means for generating comments based on emotional data" refers to a technology in which a generative AI model automatically creates comments that are appropriate to the user's emotions, taking into account collected emotional data.
[1304] The system for implementing the present invention includes a terminal, a server, a generative AI model, and an emotion engine. The specific operation of this system will be described in detail below.
[1305] The terminal is a device for users to upload video files, and can be a smartphone, tablet, or PC. Using this terminal, users upload video files they have taken. The video files are sent to the server and stored in the server's storage.
[1306] The server analyzes the uploaded video files and extracts important scenes (highlight scenes) based on the analysis results. This analysis uses a generative AI model to analyze the video content in detail. The server also converts the format of the video file as needed to make it easier to analyze.
[1307] Next, the server uses an emotion engine to sense the user's emotions in real time while watching the video and collects emotion data. This emotion data is acquired while the user is watching the video, and the user's emotional state is detected from their facial expressions, voice, gestures, etc. The data obtained by the emotion engine is sent to a generative AI model and reflected in the selection of highlight scenes and comment generation.
[1308] The generative AI model takes into account the collected emotional data and automatically generates appropriate comments for the extracted highlight scenes. These comments are in line with the video content and the user's emotions, aiming to evoke empathy in the viewer. The generated comments are added as a text overlay to the video.
[1309] The overlaid highlight footage is then converted into a format suitable for the specific platform (e.g. YouTube, Instagram, etc.) The server performs this conversion, adjusting the video resolution, aspect ratio, etc.
[1310] Finally, the server automatically distributes the converted highlight video to the specified content distribution service or social networking site, allowing users to easily publish and share the video.
[1311] As a concrete example, consider the case where a user uploads a video of a soccer match from their smartphone to a server. When the user uploads the video file of the match, the server receives the file and converts the format as necessary. The server then passes the video file to a generative AI model, which detects things like goal scenes and crowd cheers. At the same time, as the user watches, the emotion engine recognizes the user's emotions in real time and collects them as data. For example, if the user becomes excited at a goal, that emotional data is reflected in the generative AI model, which selects the scene as particularly noteworthy. The server then generates comments such as "Amazing goal!" or "Great play!"
[1312] Below are some example prompts to input to a generative AI model:
[1313] "Please extract the goal scenes and excitement of the crowd from this soccer match video, generate optimal comments based on the user's emotional data, and create a cut-out video."
[1314] In this way, the system can enhance the user's viewing experience and generate high-quality videos that evoke stronger empathy in the viewer.
[1315] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1316] Step 1:
[1317] The server receives video files uploaded by users from their devices and stores them in the server's storage. The input is the video file from the user's device, and the output is the video file stored in the server storage. In this step, file reception and storage operations are performed.
[1318] Step 2:
[1319] The server checks the metadata of the stored video file and performs format conversion if necessary. The input is the stored video file, and the output is a video file converted into a format that is easy for the generative AI model to analyze. In this step, metadata is checked and format conversion is performed.
[1320] Step 3:
[1321] The server passes the converted video file to the generative AI model, which then performs a detailed analysis of the video content. The input is the converted video file, and the output is the highlight scenes identified as a result of the video analysis. In this step, detailed video analysis is performed using the generative AI model.
[1322] Step 4:
[1323] The server uses an emotion engine to sense the user's emotions in real time while the user is watching and collects emotion data. The input is the user's facial expressions and voice while watching, and the output is the emotion data collected in real time. In this step, emotion recognition and data collection take place.
[1324] Step 5:
[1325] The server uses a generative AI model based on the collected emotion data to generate comments for the highlight scenes. The input is emotion data and the highlight scenes, and the output is the generated comments. In this step, emotion data is analyzed and comments are generated.
[1326] Step 6:
[1327] The server adds the generated comments to the video as a text overlay. The input is the highlight scene and the generated comments, and the output is the highlight video with the comments overlayed. In this step, the text overlay is added.
[1328] Step 7:
[1329] The server converts the overlaid highlight video into a specific platform format. The input is the highlight video with overlaid comments, and the output is a video format optimized for the specific platform. In this step, the format conversion takes place.
[1330] Step 8:
[1331] The server automatically distributes the converted highlight video to the specified content distribution service or social networking site. The input is a video file converted for a specific platform, and the output is a distributed and published highlight video. In this step, the video is distributed and published.
[1332] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1333] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1334] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1335] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1336] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1337] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1338] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1339] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1340] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1341] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1342] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1343] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1344] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1345] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1346] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1347] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1348] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1349] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1350] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1351] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1352] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1353] The following is further disclosed regarding the above embodiment.
[1354] (Claim 1)
[1355] How to upload video files,
[1356] A means of analyzing video files;
[1357] A means for extracting highlight scenes based on the analysis results;
[1358] a means for generating a comment for the extracted highlight scene;
[1359] A means of overlaying the generated comments onto the video;
[1360] A means to convert the overlaid highlight video into a specific platform format; and
[1361] A means to distribute the converted video to a specific platform;
[1362] A system including:
[1363] (Claim 2)
[1364] 10. The system of claim 1, further comprising means for verifying metadata of the video file and performing format conversion if necessary.
[1365] (Claim 3)
[1366] 2. The system according to claim 1, further comprising means for extracting highlight scenes based on specific keywords or scene changes during the video analysis process.
[1367] "Example 1"
[1368] (Claim 1)
[1369] A means for uploading video data;
[1370] A means for checking the metadata of video data and converting the format as necessary;
[1371] a means for analyzing the video data;
[1372] A means for extracting important scenes based on the analysis results;
[1373] a means for generating a description for the extracted important scene;
[1374] a means for overlaying the generated description onto the video;
[1375] Means of converting overlaid important videos into specific platform formats and
[1376] A means to distribute the converted video to a specific platform;
[1377] A system including:
[1378] (Claim 2)
[1379] 2. The system according to claim 1, further comprising means for extracting important scenes based on specific keywords or changes in the video during the video analysis process.
[1380] (Claim 3)
[1381] 10. The system of claim 1, further comprising means for analyzing the video data and adding automatically generated descriptive text as a text overlay for a particular video.
[1382] "Application Example 1"
[1383] (Claim 1)
[1384] A means for uploading video data;
[1385] a means for analyzing the video data;
[1386] A means for extracting important scenes based on the analysis results;
[1387] a means for generating an explanation for the extracted important scenes;
[1388] a means for overlaying the generated description onto the video;
[1389] A means for converting the superimposed important scene video into a specific distribution format;
[1390] A means for transmitting the converted video to a specific distribution platform;
[1391] A system including:
[1392] (Claim 2)
[1393] 2. The system according to claim 1, further comprising means for checking meta information of the video data and performing format conversion as necessary.
[1394] (Claim 3)
[1395] 2. The system according to claim 1, further comprising means for extracting important scenes based on specific phrases or scene changes during the video analysis process.
[1396] (Claim 4)
[1397] 2. The system according to claim 1, further comprising means for automatically extracting important scenes from the video being viewed in real time and immediately notifying and displaying them on the user's display device.
[1398] (Claim 5)
[1399] The system of claim 4 further includes a means for generating explanations for important scenes automatically extracted from a video being viewed in real time and automatically providing the explanations using a generation AI.
[1400] "Example 2: Combining Emotion Engines"
[1401] (Claim 1)
[1402] How to upload video files,
[1403] A means of analyzing video files;
[1404] A means for extracting highlight scenes based on the analysis results;
[1405] a means for generating a comment for the extracted highlight scene;
[1406] A means for recognizing and collecting user emotions;
[1407] a means for generating comments taking into account the collected sentiment data;
[1408] A means of overlaying the generated comments onto the video;
[1409] A means to convert the overlaid highlight video into a specific platform format; and
[1410] A means to distribute the converted video to a specific platform;
[1411] A system including:
[1412] (Claim 2)
[1413] 10. The system of claim 1, further comprising means for verifying metadata of the video file and performing format conversion if necessary.
[1414] (Claim 3)
[1415] 2. The system according to claim 1, further comprising means for extracting highlight scenes based on specific keywords or visual changes during the video analysis process.
[1416] "Application example 2 when combining emotion engines"
[1417] (Claim 1)
[1418] How to upload video files,
[1419] A means of analyzing video files;
[1420] A means for extracting highlight scenes based on the analysis results;
[1421] a means for generating a comment for the extracted highlight scene;
[1422] A means of overlaying the generated comments onto the video;
[1423] A means to convert the overlaid highlight video into a specific platform format; and
[1424] A means to distribute the converted video to a specific platform;
[1425] A means for sensing a user's emotions while watching in real time and collecting emotion data;
[1426] A means for generating comments based on emotion data;
[1427] A system including:
[1428] (Claim 2)
[1429] 10. The system of claim 1, further comprising means for verifying metadata of the video file and performing format conversion if necessary.
[1430] (Claim 3)
[1431] 2. The system according to claim 1, further comprising means for extracting highlight scenes based on specific keywords or scene changes during the video analysis process. [Explanation of symbols]
[1432] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. How to upload video files, A means of analyzing video files; A means for extracting highlight scenes based on the analysis results; a means for generating a comment for the extracted highlight scene; A means of overlaying the generated comments onto the video; A means to convert the overlaid highlight video into a specific platform format; and A means to distribute the converted video to a specific platform; A system including:
2. 10. The system of claim 1, further comprising means for verifying metadata of the video file and performing format conversion if necessary.
3. 2. The system according to claim 1, further comprising means for extracting highlight scenes based on specific keywords or scene changes during the video analysis process.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A