System
A generative AI model-based system efficiently detects and extracts specific scenes from videos, addressing inefficiencies in traditional methods by automating the highlight generation process.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-19
- Publication Date
- 2026-03-04
AI Technical Summary
The process of searching for specific scenes in long videos and creating highlight reels is inefficient, requiring significant time and skill, and existing systems lack accurate scene detection and highlight generation capabilities.
A system utilizing a generative AI model with natural language processing and image/video analysis to automatically detect and extract specific scenes from videos, enabling efficient generation of highlight videos.
Significantly reduces the time and skill required for video editing, allowing for rapid generation of compelling highlight content with high accuracy.
Smart Images

Figure 2026035284000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Today, video creators and everyday users spend a great deal of time and skill searching for specific scenes from long videos and creating highlight reels. This process is highly inefficient, resulting in many potentially valuable scenes being discarded. Furthermore, with the spread of online education and social media, the demand for short videos that convey information in a short amount of time is rapidly increasing, creating a need for greater efficiency in this process. However, traditional methods present challenges, such as the difficulty of accurately searching for and extracting specific scenes, which requires a great deal of effort. [Means for solving the problem]
[0005] To solve this problem, the present invention provides the following solutions. First, it provides a function for users to upload video files and input text data specifying specific scenes. Next, it uses a generative AI model to analyze the video file based on the input text data and detect specific scenes. The generative AI model includes a natural language processing model and an image / video analysis model, and identifies scenes using facial recognition and behavior recognition algorithms. It automates the process of extracting the detected scenes and generating a highlight video based on them. Furthermore, it proposes an innovative system that enables video creators and general users to efficiently search for and edit specific scenes by combining the system with a means for providing the generated highlight video to users. This invention significantly reduces the time and skill required for video editing and enables the rapid generation of more compelling content.
[0006] A "video file" is a file that records information that combines moving images and audio, and is generally used as video content.
[0007] "Upload" is the process in which a user sends data from their own device to a server and stores the data on the server.
[0008] "Text data" refers to information in the form of sentences or character strings, and in the present invention, in particular, refers to keywords and explanatory text for specifying a particular scene.
[0009] A "generative AI model" is an algorithm or set of algorithms that uses artificial intelligence techniques to analyze data and generate results based on a specific purpose.
[0010] A "natural language processing model" is an AI model designed to understand, analyze, and generate human language, and is used to parse the meaning of text data.
[0011] An "image and video analysis model" is an AI model that analyzes image and video data and recognizes specific objects and actions.
[0012] A "specific scene" refers to a portion of a long video in which a specific event or phenomenon occurs.
[0013] "Detection" is the process by which a generative AI model finds the desired scene based on specified features in a video.
[0014] "Extraction" is the process of cutting out the detected scenes from the video file and extracting them as individual clips.
[0015] A "highlight video" is a clip that is created by selecting and compiling specific scenes from a video and then editing them into a clip that can be viewed in a short amount of time.
[0016] "Providing" refers to the process of releasing the generated highlight video to the public in a form that allows users to view or download it.
[0017] "User" refers to an individual or organization that uses this system, and is the entity that uploads video files and inputs text data.
[0018] A "server" is a computer system that provides services and data over a network, and in the present invention stores, analyzes, and edits video files.
[0019] A "terminal" is a device that is directly operated by a user, and is used to upload video files and input text data. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0022] First, the terms used in the following description will be explained.
[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0028] [First embodiment]
[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0041] The present invention is applied to a system that automatically searches for specific scenes from a video file and generates a highlight video. The system mainly includes a user, a terminal, and a server.
[0042] Basic system configuration
[0043] The system according to the present invention comprises the following main components:
[0044] 1. User's device: A device where video files are uploaded and text data specifying specific scenes is entered.
[0045] 2. Server: A computer system that uses a generative AI model to analyze video and extract scenes, then generates a highlight video.
[0046] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models to search and extract specific scenes.
[0047] Program processing flow
[0048] Upload video and enter text
[0049] Users upload video files to the server through a browser or a dedicated application. They also enter keywords and descriptions to specify specific scenes in a text input field. The device sends this data to the server as a single HTTP request.
[0050] Receiving and initial processing of input data
[0051] The server analyzes the received HTTP request and extracts the video file and text data. The video file is temporarily stored, and the text data is used for the next stage of analysis.
[0052] Launching and analyzing generative AI models
[0053] The server launches the generative AI model. First, text data is input into the natural language processing model, and keyword and semantic analysis is performed. As a specific example, the keyword "Yanagida's turn at bat" is analyzed and its features are generated.
[0054] Next, based on the generated features, the image and video analysis model scans the entire video to search for relevant scenes, and uses facial and behavioral recognition algorithms to identify the start and end times of scenes that match the analyzed features.
[0055] Scene detection and extraction
[0056] The server detects scenes based on the output data from the generative AI model. Using the time information of the identified scenes, it extracts the scenes using a video editing library (e.g., FFmpeg or OpenCV). The extracted scenes are temporarily saved as individual clips.
[0057] Highlight video generation and provision
[0058] The server combines the extracted scenes and edits them into a highlight video. The completed highlight video is saved and a URL for the saved location is generated. This URL is provided to the user, who can then view or download the generated highlight video via the link.
[0059] Specific examples
[0060] For example, if a user is searching for a scene of "Yanagida's turn at bat in a baseball broadcast in 2023," the process would be as follows: First, the user uses their device to upload a video file of the baseball broadcast to the server and enters the keyword "Yanagida's turn at bat." The server inputs the text data and video file into the generative AI model, and performs keyword analysis using a natural language processing model.
[0061] Next, an image and video analysis model scans the entire video to search for and detect scenes that correspond to Yanagi's at-bat. The identified scenes are extracted and edited to create a highlight video. Finally, users can access the generated highlight video using the provided link.
[0062] In this way, the system according to the present invention enables the user to easily search for a particular scene and realizes the process of efficiently generating a highlight video.
[0063] The processing flow will be explained below.
[0064] Step 1:
[0065] Users select video files from their own devices and upload them to the server via a browser or dedicated application.
[0066] Step 2:
[0067] The user inputs text data specifying a particular scene into the input field, for example, specifying the keyword "Yanagida's turn at bat."
[0068] Step 3:
[0069] The terminal sends the video file and text data to the server as a single HTTP request.
[0070] Step 4:
[0071] The server analyzes the received HTTP request and extracts the video file and text data. The video file is temporarily stored, and the text data is used to analyze the generative AI model.
[0072] Step 5:
[0073] The server launches the generative AI model. First, it inputs text data into the natural language processing model, analyzes the keyword "Yanagida's turn at bat," and generates its features.
[0074] Step 6:
[0075] The server inputs the generated features into an image / video analysis model, which then scans the entire video to find scenes that match the features.
[0076] Step 7:
[0077] The image and video analysis model detects the start and end times of scenes corresponding to "Yanagita's turn at bat" from the video and returns them to the server.
[0078] Step 8:
[0079] The server extracts the scene from the video file using a video editing library (e.g., FFmpeg or OpenCV) based on the time information of the detected scene.
[0080] Step 9:
[0081] The server temporarily stores the extracted scenes as one clip.
[0082] Step 10:
[0083] The server combines a plurality of extracted scenes as necessary and edits them into a highlight video.
[0084] Step 11:
[0085] The server saves the completed highlight video and generates a URL for the location where it is saved.
[0086] Step 12:
[0087] The server sends the generated URL to the terminal as an HTTP response.
[0088] Step 13:
[0089] The terminal displays the received URL to the user.
[0090] Step 14:
[0091] The user can click on the displayed link to view or download the generated highlight video.
[0092] Example 1
[0093] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0094] In modern video media, the task of efficiently extracting specific scenes from long video data and generating the highlight footage desired by users is extremely time-consuming and labor-intensive. Manually editing large amounts of video data is particularly impractical, creating a demand for automated systems. However, existing systems lack the technology for highly accurate scene detection and highlight generation. To address this issue, a system that achieves more accurate scene detection and highlight generation is needed.
[0095] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0096] In this invention, the server includes means for uploading video data, means for inputting text data specifying specific scenes, means for analyzing the video data based on the input text data using a generative AI model to detect the specific scenes, means for extracting the detected scenes and generating a highlight video, and means for providing the generated highlight video to the user. This makes it possible to automatically extract scenes specified by the user with high accuracy and easily generate a highlight video.
[0097] "Video data" refers to video files stored in digital format and has the attributes of a media file.
[0098] "Text data" is character string information that the user inputs to specify a particular scene.
[0099] A "generative AI model" is a model that uses artificial intelligence technology and includes a series of algorithms to analyze video data based on text data and detect specific scenes.
[0100] A "natural language processing model" is part of a generative AI model and is an algorithm for analyzing the meaning of input text data and extracting keywords.
[0101] The "image and video analysis model" is part of the generative AI model and is an algorithm for searching and detecting specific scenes within video data.
[0102] A "person recognition algorithm" is a technology for identifying specific people within video data.
[0103] "Movement recognition algorithms" are techniques for identifying specific movements or actions within video data.
[0104] "Cutting out" is an operation for extracting a specific portion of video data based on the start and end times of a detected scene.
[0105] A "highlight video" is a short video file created by combining specific scenes, and is a format that allows users to efficiently view scenes they want to focus on.
[0106] A "user" is an individual or group that operates the system to upload video data and specify specific scenes.
[0107] The present invention is applied to a system that automatically searches for specific scenes from video data and generates highlight videos. The system mainly includes a user, a terminal, and a server.
[0108] Basic system configuration
[0109] The system according to the present invention comprises the following main components:
[0110] 1. Terminal: A device where a user uploads video data and inputs text data specifying a particular scene.
[0111] 2. Server: A computer system that uses a generative AI model to analyze video data, extract scenes, and generate highlight footage.
[0112] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models that search for and extract specific scenes.
[0113] Example of operation
[0114] For example, if a user requests "a specific player's play from a sports broadcast in 2023," the process would proceed as follows: The user uses their device to upload a video file of the sports broadcast to the server and enters the keyword "a specific player's play." The server inputs the text data and video data into the generative AI model, and performs keyword analysis using a natural language processing model.
[0115] Next, based on the generated features, the image and video analysis model scans the entire video data to search for relevant scenes, and uses facial and behavior recognition algorithms to identify the start and end times of scenes that match the analyzed features.
[0116] Once the scenes are identified, the server uses a video editing library (e.g., FFmpeg or OpenCV) to extract the segments, which are then temporarily saved as individual clips.
[0117] Finally, the server combines the extracted scenes and edits them into a highlight video. The completed highlight video is saved on the server, and a URL for the saved location is generated. This URL is provided to users, who can then view or download the created highlight video via the link.
[0118] Hardware and software used
[0119] Hardware: User's device (PC, smartphone, etc.), server
[0120] Software: Generative AI models, video editing libraries (FFmpeg and OpenCV), natural language processing models (BERT, GPT, etc.), image and video analysis models (YOLO, etc.)
[0121] Prompt Sentence Examples
[0122] "Automatically generate highlight footage including a specific player's play from a sports broadcast in 2023."
[0123] By using this system, users can efficiently extract specific scenes from long video data and generate highlight videos in a short amount of time.
[0124] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0125] Step 1:
[0126] The user uploads video data using a device and enters text data to specify a specific scene. Specifically, the user selects video data through a browser or a dedicated application and enters keywords such as "a specific player's play." The device then sends this data to the server as a single HTTP request.
[0127] Input: Video data, text data
[0128] Output: HTTP request
[0129] Step 2:
[0130] The server analyzes the received HTTP request and extracts video and text data. The server temporarily stores the video data, and the text data is used in the next analysis stage. A parsing process is performed to analyze and extract data from the HTTP request. Keywords such as "plays by a specific player" are extracted and stored as input data.
[0131] Input: HTTP request
[0132] Output: Video data, text data
[0133] Step 3:
[0134] The server launches the generative AI model. First, the text data is input into the natural language processing model, which performs keyword analysis and semantic analysis. As a specific example, the natural language processing model analyzes the "play of a specific player" and generates its features. These features become indicators for searching the video data.
[0135] Input: Text data
[0136] Output: Text features
[0137] Step 4:
[0138] The server launches the image and video analysis model, scans the video data based on the generated text features, and searches for relevant scenes. It then uses facial and behavior recognition algorithms to identify the start and end times of specific scenes. Specifically, it analyzes the video data frame by frame to find scenes that match the features.
[0139] Input: Video data, text features
[0140] Output: Scene start and end times
[0141] Step 5:
[0142] The server uses a video editing library (e.g., FFmpeg, OpenCV) to extract the identified scene based on its time information. The extracted scene is temporarily saved as an individual clip. Specifically, the server uses a library such as FFmpeg to extract video within a specified time range.
[0143] Input: Video data, scene start and end times
[0144] Output: Video clip
[0145] Step 6:
[0146] The server combines multiple scenes and edits them into a highlight video. The completed highlight video is saved on the server and a URL for its location is generated. Specifically, an editing library such as FFmpeg is used to combine the individual clips into a single video file.
[0147] Input: Video clip
[0148] Output: highlight video file, viewing URL
[0149] Step 7:
[0150] The user can view or download the generated highlight video using the provided viewing URL. Specifically, the user opens the URL in a browser and plays or saves the video.
[0151] Input: Viewing URL
[0152] Output: Played video or downloaded video file
[0153] (Application example 1)
[0154] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0155] Conventional video search systems have had the problem that it takes a lot of time and effort to find a specific scene. Manually editing a highlight video is also cumbersome, placing a significant burden on the user. Furthermore, there has been a lack of means for efficiently viewing the generated highlight video, making it difficult to improve the user experience. The present invention aims to solve these problems by providing a system that enables users to easily search for specific scenes, generate a highlight video, and efficiently view it.
[0156] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0157] In this invention, the server includes means for uploading video files, means for inputting text data specifying specific scenes, means for analyzing the video files based on the input text data using a generative AI model and detecting specific scenes, means for extracting the detected scenes and generating a highlight video, means for using cloud storage to provide the generated highlight video to a user, and means for providing the generated highlight video by streaming, thereby enabling a user to efficiently search for specific scenes with little effort, generate a highlight video, and further watch the highlight video by streaming.
[0158] "Means for uploading video files" refers to a function that allows a user to send video data from their own device to a server and save it.
[0159] "Means for inputting text data specifying a particular scene" refers to an input interface that a user uses to specify a scene to be searched for using keywords or a description.
[0160] "Means of using a generative AI model to analyze video files based on input text data and detect specific scenes" refers to the function in which a generative AI model analyzes text data and identifies the relevant scenes in the video based on the results.
[0161] The "means for extracting detected scenes and generating a highlight video" refers to a function for cutting out identified scenes from the original video and editing them to create a highlight video.
[0162] The "means of using cloud storage to provide the generated highlight video to the user" refers to a function for storing the generated highlight video on the cloud and enabling the user to access it.
[0163] The term "means for providing the generated highlight video by streaming" refers to a function of transmitting the video via the Internet so that the user can view the highlight video generated in real time.
[0164] Basic system configuration
[0165] The system according to the present invention comprises the following main components:
[0166] 1. User's device: A device on which a video file is uploaded using a smartphone app and text data specifying a specific scene is entered.
[0167] 2. Server: A computer system that performs video analysis and scene extraction, and provides the generated highlight video to users. This server runs on Amazon AWS (registered trademark) or other cloud platforms.
[0168] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models that search for and extract specific scenes.
[0169] Upload video and enter text
[0170] A user uploads a video file to a server through a smartphone app. Then, they enter keywords and descriptions to specify specific scenes in a text input field. For example, they enter the keyword "live performance of a popular artist." This input data is sent from the application to the server as a single HTTP request.
[0171] Receiving and initial processing of input data
[0172] The server analyzes the received HTTP request and extracts the video file and text data. The video file is temporarily stored in cloud storage such as Amazon S3, and the text data is used for the next stage of analysis.
[0173] Launching and analyzing generative AI models
[0174] The server launches a generative AI model. First, it inputs text data into a natural language processing (NLP) model (e.g., GPT-4 (registered trademark)) to perform keyword and semantic analysis. This operation identifies the content that the input keywords contain. Next, it uses an image / video analysis model (e.g., OpenCV or YOLO) to scan the entire video based on the analyzed keywords and search for relevant scenes.
[0175] Scene detection and extraction
[0176] Using the time information of the identified scene, the part is extracted using a video editing library (e.g., FFmpeg). The detected scene is temporarily saved as a clip.
[0177] Highlight video generation and provision
[0178] The server combines multiple scenes and edits them into a highlight video. The generated highlight video is saved to a server such as Amazon S3, and a URL for the saved location is generated. Users can view or download the generated highlight video from the app via this URL. The highlight video is also provided in real time via a streaming service.
[0179] Specific examples
[0180] For example, if a user searches for "live performances by popular artists," the process would proceed as follows: The user uses their device to upload a video file of the live performance to the server and enters the keyword "live performances by popular artists." The server uses a generative AI model to analyze the text data and video file, and a natural language processing model to perform keyword analysis. Next, an image and video analysis model scans the entire video to search for relevant scenes, and identified scenes are extracted as individual clips. These clips are then edited to create a highlight video, which is then provided to the user.
[0181] Prompt Sentence Examples
[0182] "Search for live performance scenes of popular artists and generate highlight videos."
[0183] In this way, the system according to the present invention realizes a process that allows a user to easily search for a particular scene and efficiently generate and view a highlight video.
[0184] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0185] Step 1:
[0186] A user uploads a video file to a server through a smartphone app. The input is a video file from the user's device, and the output is a video file stored on the server. Specifically, the user uses the app's "video upload" function to send a video file via an HTTP request. The server receives the file and stores it in cloud storage such as Amazon S3.
[0187] Step 2:
[0188] Within the same application, the user enters text data to specify a particular scene. The input is keywords and descriptions related to the particular scene, and the output is text data sent to the server. Specifically, the user enters keywords in the text field and presses the "Search" button. The text data is sent to the server along with the video file via an HTTP request.
[0189] Step 3:
[0190] The server analyzes the received video file and text data. The input is the video file and text data, and the output is the text analysis results. Specifically, the server retrieves the video file from Amazon S3 and simultaneously supplies the text data to a natural language processing (NLP) model (e.g., GPT-4). The NLP model performs keyword analysis and semantic analysis to generate the necessary features.
[0191] Step 4:
[0192] The image and video analysis part of the generative AI model is activated, and it scans the entire video based on the analyzed features to search for specific scenes. The input is the analysis results from the NLP model and the video file, and the output is the time information of the detected scene. Specifically, the generative AI model (e.g., OpenCV or YOLO) analyzes the entire video frame by frame and identifies the start and end times of the relevant scenes.
[0193] Step 5:
[0194] The server extracts the detected scenes and generates highlight clips. The input is the scene's time information and the original video file, and the output is the extracted highlight clip. Specifically, the server uses a video editing library (e.g., FFmpeg) to extract a specific scene from its start time to its end time and temporarily save it as an individual clip.
[0195] Step 6:
[0196] The server then combines the extracted scenes into a single highlight video. The input is multiple highlight clips, and the output is the completed highlight video. Specifically, the server uses FFmpeg again to combine the temporarily saved clips into a single video file.
[0197] Step 7:
[0198] The server saves the generated highlight video in cloud storage and generates a URL for the storage location. The input is the completed highlight video, and the output is the video URL. Specifically, the server uploads the completed highlight video to Amazon S3 and generates an access URL.
[0199] Step 8:
[0200] The user uses the provided URL to watch or download the generated highlight video. The input is the video URL, and the output is the highlight video played on the user's device. Specifically, the user clicks the URL link in the smartphone app and streams or downloads the video for viewing.
[0201] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0202] This invention is applied to a system that automatically searches for specific scenes from video files and generates highlight videos by combining an emotion engine that recognizes the user's emotions. The system selects scenes that correspond to the user's emotional state.
[0203] Basic system configuration
[0204] The system according to the present invention comprises the following main components:
[0205] 1. User's device: A device where video files are uploaded and text data specifying specific scenes is entered.
[0206] 2. Server: A computer system that uses a generative AI model to analyze videos and extract scenes, and combines it with the user's emotion engine to generate a highlight video.
[0207] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models to search and extract specific scenes.
[0208] 4. Emotion Engine: An algorithm to recognize the user's emotions and search and select scenes according to their emotional state.
[0209] Program processing flow
[0210] Upload video and enter text
[0211] Users upload video files to the server through a browser or a dedicated application. They also enter keywords or descriptions to identify specific scenes in a text input field. The system also collects data to recognize the user's emotions. For example, it records facial expression data and voice input while the user is watching a specific video scene.
[0212] Receiving and initial processing of input data
[0213] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, and the text data and emotion data are used for analysis by the generative AI model and emotion engine.
[0214] Launching and analyzing generative AI models
[0215] The server launches the generative AI model. First, text data is input into the natural language processing model, and keyword and semantic analysis is performed. As a specific example, the keyword "Yanagida's turn at bat" is analyzed and its features are generated.
[0216] Next, based on the generated features, the image and video analysis model scans the entire video to find relevant scenes. In addition, the emotion engine analyzes the user's emotional data and determines their emotional state.
[0217] Scene detection and extraction
[0218] The server detects scenes based on the output data from the generative AI model and the emotion analysis results from the emotion engine. For example, if the emotion engine recognizes a high level of happiness in the user, it determines that the scene is particularly important and sets a high priority for it.
[0219] Highlight video generation and provision
[0220] The server uses the time information of the detected scenes to extract those parts using a video editing library (e.g., FFmpeg or OpenCV). The extracted scenes are temporarily saved as a single clip. Multiple extracted scenes are then combined as needed to create a highlight video. The completed highlight video is saved, and a URL for its save location is generated.
[0221] Reflecting user reactions based on emotional data
[0222] The system accumulates user emotion data through the emotion engine, which can be used to improve the accuracy of future analysis results. Also, by reflecting the user's emotions toward specific scenes, the system can provide more accurate highlight videos.
[0223] Specific examples
[0224] For example, if a user is searching for a scene of "Yanagida's turn at bat in a baseball broadcast in 2023," the process would be as follows. First, the user uses their device to upload a video file of the baseball broadcast to the server and enters the keyword "Yanagida's turn at bat." The emotion engine then collects facial expressions and voice data while the user is watching the video. The server then inputs the text data and video file into a generative AI model, and performs keyword analysis using a natural language processing model.
[0225] Next, an image and video analysis model scans the entire video to search for and detect scenes that correspond to Yanagi's turn at bat. At the same time, an emotion engine analyzes the user's emotional data and prioritizes scenes in which the user expresses happiness. The identified scenes are extracted and edited to create a highlight video. Finally, the user can access the generated highlight video using the provided link.
[0226] In this way, by combining emotion engines, the present invention enables scene selection according to the individual emotional state of the user, and generates attractive content with high accuracy.
[0227] The processing flow will be explained below.
[0228] Step 1:
[0229] Users select video files from their own devices and upload them to the server via a browser or dedicated application.
[0230] Step 2:
[0231] The user inputs text data specifying a particular scene into the input field, for example, specifying the keyword "Yanagida's turn at bat."
[0232] Step 3:
[0233] The user gives permission to the system to record emotional data (facial expressions and voice) while watching the video.
[0234] Step 4:
[0235] The terminal transmits the video file, text data, and emotion data to the server as a single HTTP request.
[0236] Step 5:
[0237] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, and the text data and emotion data are used for analysis by the generative AI model and emotion engine.
[0238] Step 6:
[0239] The server launches the generative AI model. Text data is input into the natural language processing model, and semantic analysis of the text is performed. For example, the keyword "Yanagida's turn at bat" is analyzed and features are generated.
[0240] Step 7:
[0241] The server runs an image / video analysis model based on the generated features and searches the entire video for corresponding scenes. The image / video analysis model detects specific scenes using facial and behavior recognition algorithms.
[0242] Step 8:
[0243] The emotion engine analyzes the user's emotion data (facial expression data and voice data) to determine the user's emotional state, for example, identifying emotion categories such as happiness, excitement, and anxiety.
[0244] Step 9:
[0245] Based on the analysis results of the emotion engine, the server evaluates scenes in the video based on the emotion data and feature values, and sets priorities. Scenes in which the user expressed a favorable emotion are given a higher priority.
[0246] Step 10:
[0247] Based on the evaluation results, the server extracts specific scenes from the video file by identifying the start and end times of the scenes using a video editing library (such as FFmpeg or OpenCV) and extracting those parts.
[0248] Step 11:
[0249] The server temporarily stores the extracted scenes as one clip.
[0250] Step 12:
[0251] The server combines multiple extracted scenes as needed and edits them into a highlight video based on the user's emotional state.
[0252] Step 13:
[0253] The server saves the completed highlight video and generates a URL for the location where it is saved.
[0254] Step 14:
[0255] The server sends the generated URL to the terminal as an HTTP response.
[0256] Step 15:
[0257] The terminal displays the received URL to the user.
[0258] Step 16:
[0259] The user can click on the displayed link to view or download the generated highlight video.
[0260] Example 2
[0261] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0262] Conventional systems were unable to consider the user's subjective emotions when searching for specific scenes within a video file and generating a highlight video. This made it difficult to appropriately select the scenes that appealed to each individual user. Furthermore, there was a lack of technology for performing highly accurate analysis using user emotion data. This has led to a demand for improved user satisfaction.
[0263] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0264] In this invention, the server includes means for uploading video files, means for inputting text data specifying specific scenes, means for analyzing the video files based on the input text data using a generative AI model and detecting specific scenes, means for collecting user emotion data, means for analyzing the emotion data and selecting scenes based on the user's emotional state, means for extracting the detected scenes and generating a highlight video, and means for providing the generated highlight video to the user. This enables scene selection that takes user emotion into consideration and highly accurate highlight video generation.
[0265] 1. "Means for uploading video files" means the technical means that allows a user to transmit video files to a server via the Internet.
[0266] 2. "Means for inputting text data specifying a specific scene" means a technical means for providing an interface for a user to input text data such as keywords or descriptions in order to specify a specific video scene.
[0267] 3. "Generative AI model" refers to a set of algorithms that analyze input data and process it for a specific purpose, including, specifically, natural language processing models and image and video analysis models.
[0268] 4. "Natural language processing model" refers to an algorithm for analyzing text data and performing keyword analysis and semantic analysis.
[0269] 5. "Image / Video Analysis Model" refers to an algorithm that analyzes scenes in a video file and detects specific actions or objects.
[0270] 6. "Means for collecting user emotional data" refers to technical means for detecting a user's facial expressions, voice, actions, etc., and collecting emotional data based on them.
[0271] 7. "Means for analyzing emotional data and selecting scenes based on the user's emotional state" means technical means for determining the user's emotional state using collected emotional data and selecting important scenes based on that emotional state.
[0272] 8. "Means for extracting detected scenes and generating a highlight video" refers to the technical means of using a video editing library to extract that portion based on the time information of the detected scenes and edit it into a single highlight video.
[0273] 9. "Means for providing the generated highlight video to the user" refers to the technical means for saving the completed highlight video, generating a URL for the location where the video is saved, and providing the URL to the user.
[0274] In this way, by combining each of these means, it is possible to generate a highlight video that reflects the user's emotions.
[0275] The present invention relates to a system that automatically searches for specific scenes from video files and generates highlight videos by combining an emotion engine that recognizes the user's emotions. This system selects scenes that correspond to the user's emotional state.
[0276] Basic system configuration
[0277] The system according to the present invention comprises the following main components:
[0278] 1. On the user's device:
[0279] A user uses a device (such as a computer or smartphone) to upload a video file.
[0280] The device sends the video file to the server via a browser or a dedicated application.
[0281] A text entry field is used to enter keywords or descriptions to specify a particular scene.
[0282] It is also equipped with sensors to collect user emotional data, such as facial expression data and voice data.
[0283] 2. Server:
[0284] The server uses a generative AI model to analyze the video and extract scenes.
[0285] It analyzes received HTTP requests and extracts and saves video files, text data, and emotional data.
[0286] The hardware used is a computer system equipped with a high-performance processor and sufficient memory.
[0287] 3. Generative AI Model:
[0288] It consists of a set of algorithms including natural language processing models and image and video analysis models.
[0289] The natural language processing model analyzes the input text data and performs keyword analysis and semantic analysis.
[0290] The image and video analysis model uses the generated features to scan the entire video and search for relevant scenes.
[0291] 4. Emotion Engine:
[0292] An algorithm for recognizing the user's emotions and searching and selecting scenes according to their emotional state.
[0293] Facial expression and voice data are analyzed to evaluate the user's emotional state.
[0294] Example of operation
[0295] For example, if a user searches for a scene of a specific player playing in a sports broadcast in 2023, they would proceed as follows:
[0296] 1. Upload video and enter text:
[0297] A user uses his / her own terminal to upload a video file of a live sports broadcast to a server.
[0298] Enter a keyword, such as "a specific player's play," into the text entry field.
[0299] While the user is watching the video, facial expression and voice data is collected by the device's sensors.
[0300] 2. Receiving input data and initial processing:
[0301] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data.
[0302] The video file is temporarily saved, and text and emotion data are passed to the generative AI model and emotion engine.
[0303] 3. Launch and analyze the generative AI model:
[0304] The server launches the generative AI model, and the natural language processing model analyzes the text data. It performs keyword analysis based on the keyword "play by a specific player."
[0305] The image and video analysis model scans the entire video based on the generated features and searches for relevant scenes.
[0306] The emotion engine analyzes facial expression and voice data to determine the user's emotional state.
[0307] 4. Scene detection and extraction:
[0308] The server detects scenes based on the output data from the generative AI model and the emotion analysis results from the emotion engine.
[0309] If the emotion engine rates the user's happiness highly, it sets the priority of that scene high.
[0310] 5. Highlight video generation and delivery:
[0311] The server extracts the detected scene using a video editing library (e.g., FFmpeg or OpenCV) based on the time information of the detected scene.
[0312] Multiple extracted scenes are combined and edited into a highlight video.
[0313] The completed highlight video is saved, and a URL for the saved location is generated and provided to the user.
[0314] In this way, by combining emotion engines, the present invention makes it possible to select scenes according to the individual emotional state of the user, and to generate attractive content with high accuracy.
[0315] Prompt Sentence Examples
[0316] It searches for and extracts specific scenes based on text data entered by the user, such as "a specific player's play from a sports broadcast in 2023."
[0317] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0318] Step 1: Upload video and enter text
[0319] Users select video files from a dedicated application or browser on their own devices and upload them to the server. Users enter keywords or descriptions in a text input field to specify specific scenes. Emotional data such as facial expressions and voice data are also collected while the user is watching the video.
[0320] Input: Video file, text data (e.g., "A specific player's play"), emotional data (facial expressions, voice)
[0321] Output: The video file, text data, and emotion data sent by the user are uploaded to the server.
[0322] Specific operation: The user selects a video file and clicks the upload button. Next, the user enters a keyword specifying a specific scene in the text input field and clicks the send button. While watching the video, the user's facial expression data and voice data are collected by the device's sensors.
[0323] Step 2: Receiving input data and initial processing
[0324] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, and the text data and emotion data are prepared for analysis by the generative AI model and emotion engine.
[0325] Input: Video files uploaded to the server, text data, emotion data
[0326] Output: Temporarily saved video files, analyzable text data and emotion data
[0327] Specific operation: The server analyzes the received HTTP request and saves the video file in the specified directory. Next, it stores the text data and emotion data in memory and prepares for analysis.
[0328] Step 3: Launch and analyze the generative AI model
[0329] The server launches the generative AI model. Text data is input into the natural language processing model, which performs keyword and semantic analysis. Based on the analyzed features, the image and video analysis model scans the entire video to search for relevant scenes. At the same time, the emotion engine analyzes the user's emotional data and determines their emotional state.
[0330] Input: text data, video files, emotion data
[0331] Output: Keyword analysis results, candidate scenes, and user's emotional state
[0332] How it works: The server analyzes the text data entered into the natural language processing model and analyzes the keyword "a specific player's play." Next, based on the generated features, the image and video analysis model scans the entire video to search for specific scenes. At the same time, the emotion engine analyzes facial expression and voice data to evaluate the user's emotions.
[0333] Step 4: Scene detection and extraction
[0334] The server detects scenes based on the output data of the generative AI model and the emotion analysis results of the emotion engine. The emotion engine selects scenes that evoke a specific emotion (e.g., happiness) and sets a priority.
[0335] Input: Analysis results of the generative AI model, analysis results of the emotion engine
[0336] Output: Time information of detected scenes, prioritized scenes
[0337] Specific operation: The server identifies the start and end times of scenes corresponding to Yanagida's at-bats, sets priorities based on the results of emotion analysis, and adds high-priority scenes to a list.
[0338] Step 5: Generate and serve highlight videos
[0339] The server uses a video editing library (e.g., FFmpeg or OpenCV) to extract the relevant parts based on the time information of the detected scenes. If necessary, multiple scenes are combined and edited into a highlight video. The completed highlight video is saved, and a URL for the saved location is generated and provided to the user.
[0340] Input: Time information of detected scenes, prioritized scenes
[0341] Output: Saved highlight video, access URL to the highlight video
[0342] Specific operation: The server uses FFmpeg commands to extract specific scenes from a video file, combines multiple scenes into a single highlight video, and saves the generated highlight video to storage. Finally, it generates a URL for the highlight video and notifies the user.
[0343] (Application example 2)
[0344] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0345] In conventional highlight video generation systems, the detection and extraction of specific scenes is often performed solely through analysis of text data, without taking into account the user's emotions or preferences, resulting in the generated highlight videos not necessarily meeting the user's expectations. Furthermore, personalized content that reflects the user's emotions is in demand, particularly in entertainment applications, but few systems have this functionality.
[0346] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0347] In this invention, the server includes means for uploading video files, means for inputting text data specifying specific scenes, means for analyzing the video files based on the input text data using a generative AI model to detect specific scenes, means for extracting the detected scenes and generating a highlight video, means for collecting user emotion data, means for determining the importance of scenes based on the emotion data, and means for providing the generated highlight video to the user, thereby enabling the generation of a personalized highlight video that takes user emotions into consideration.
[0348] A "video file" is a digital file containing visual and audio information.
[0349] "Uploading" is the act of transferring data from a user's terminal to a remote computer such as a server.
[0350] "Text data" is digital data that contains written information.
[0351] A "generative AI model" is a set of artificial intelligence algorithms designed to perform natural language processing and image and video analysis.
[0352] An "emotion engine" is an algorithm that analyzes the user's emotional state and performs processing based on that data.
[0353] A "highlight video" is a shortened version of a video that has been edited to extract specific scenes.
[0354] "User" means a person or organization that uses the system.
[0355] A "natural language processing model" is an artificial intelligence model for understanding and analyzing text data.
[0356] An "image and video analysis model" is an artificial intelligence model for analyzing and recognizing the content of images and videos.
[0357] "Facial recognition" is a technology that detects and identifies human faces from images and videos.
[0358] "Behavior recognition" is a technology that detects and identifies human actions and movements from images and videos.
[0359] "Emotion data" is digital data that expresses the user's emotional state using numerical values and categories.
[0360] "Decision making" is the process of reaching a conclusion based on data.
[0361] System Overview
[0362] This invention relates to a system that recognizes a user's emotions, automatically searches and extracts specific scenes from videos based on those emotions, and generates highlight videos. The system consists of the following main components:
[0363] 1. Terminal: A device where a user uploads video files and enters text data specifying specific scenes. Typically, this is a smartphone or tablet.
[0364] 2. Server: A computer system that uses a generative AI model to analyze video files, detect and extract specific scenes, and generate a highlight video.
[0365] 3. Generative AI model: A set of algorithms, including natural language processing models and image and video analysis models, used to search and extract specific scenes.
[0366] 4. Emotion engine: An algorithm that analyzes the user's emotional data and determines the importance of a scene based on that data.
[0367] System Operation
[0368] Upload video and enter keywords
[0369] Users upload video files to the server via a smartphone app. They input keywords and descriptions in text format to specify specific scenes. While the user is watching the video, the camera and microphone are used to collect facial expressions and voice data.
[0370] Data reception and initial processing
[0371] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily saved, and the text data is input into a natural language processing model. The emotion data is sent to the emotion engine.
[0372] Analysis of generative AI models
[0373] The server launches a generative AI model to analyze the text data. A natural language processing model (e.g., the TENSORFLOW® model) performs keyword and semantic analysis and generates feature vectors. Next, an image and video analysis model scans the entire video based on the generated feature vectors to search for relevant scenes. At the same time, an emotion engine analyzes the user's emotional data and determines their emotional state (e.g., the Microsoft® Azure® Emotion API).
[0374] Scene detection and extraction
[0375] Based on the analysis results of the generative AI model and emotion engine, the server detects specific scenes. It determines the importance of each scene based on the emotion data and extracts high-priority scenes. It then uses a video editing library (e.g., FFmpeg) to extract the detected scenes and generate a highlight video.
[0376] Highlight video provided
[0377] The generated highlight video is stored on the server and a generated link is provided to the user, who can use the link to view and share the highlight video.
[0378] Specific examples
[0379] For example, if a user wants to search for "scenes of a specific player from a sports broadcast in 2023," the process would proceed as follows: The user uploads a video file of the sports broadcast to the server using their device and enters the keyword "scenes of a specific player." The emotion engine also collects facial expressions and voice data while the user is watching the video. The server inputs the text data and video file into the generative AI model, and performs keyword analysis using a natural language processing model. The image and video analysis model then scans the entire video to search for and detect relevant scenes. At the same time, the emotion engine analyzes the user's emotional data and selects scenes with high importance. The identified scenes are then extracted and edited to create a highlight video. Finally, the user can access the highlight video using the generated link.
[0380] Examples of prompt statements
[0381] "Look for scenes that include a specific player."
[0382] "Create important highlights based on the emotional data of this scene."
[0383] This system enables the generation of personalized highlight videos that reflect the user's emotions.
[0384] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0385] Step 1:
[0386] A user uploads a video file using a smartphone app and inputs text data (keywords and descriptions) to specify a specific scene. While the user is watching the video, the device uses a camera and microphone to collect facial expressions and voice data. The collected data is then sent to a server.
[0387] Input: Video files, text data, facial expressions and audio data
[0388] Output: Send data to the server
[0389] Specific behavior:
[0390] The application accepts user operations and displays a UI that allows the user to select a video file.
[0391] Provide a text field for entering keywords and descriptions.
[0392] It asks for permission to use the camera and microphone, and if the user agrees, it begins collecting data.
[0393] Step 2:
[0394] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, the text data is sent to a natural language processing model (e.g., TensorFlow model), and the emotion data is sent to an emotion engine (e.g., Microsoft Azure Emotion API).
[0395] Input: HTTP request
[0396] Output: Video file, text data, emotion data extraction
[0397] Specific behavior:
[0398] The server receives the HTTP request and parses the video file and text data.
[0399] Temporarily save video files to a file system or cloud storage.
[0400] The text data and emotion data are sent to the corresponding analysis units.
[0401] Step 3:
[0402] The generative AI model analyzes the text data. The natural language processing model performs keyword and semantic analysis and generates feature vectors. Based on these feature vectors, the image and video analysis model scans the entire video and searches for relevant scenes.
[0403] Input: Text data
[0404] Output: Features, list of detected scenes
[0405] Specific behavior:
[0406] A natural language processing model receives the text data and analyzes it for keywords and meaning.
[0407] Features are generated and the data is passed to an image / video analysis model.
[0408] The model scans the video file and lists the start and end times of the relevant scenes.
[0409] Step 4:
[0410] The emotion engine analyzes the user's emotion data and determines the emotional state the user is in. Based on the analysis results, the importance of the scene is determined.
[0411] Input: Emotion data
[0412] Output: Emotional state analysis result, scene importance judgment
[0413] Specific behavior:
[0414] The emotion engine analyzes the user's facial expressions and voice data to detect emotional states such as happiness, surprise, and excitement.
[0415] Emotional data is collected for each scene and the importance of that scene is calculated.
[0416] Step 5:
[0417] The server detects and extracts specific scenes based on the analysis results of the generated AI model and emotion engine, and uses a video editing library (e.g., FFmpeg) to extract the detected scenes and generate a highlight video.
[0418] Input: List of scenes, emotional state analysis results
[0419] Output: Highlight video
[0420] Specific behavior:
[0421] The list of detected scenes is compared with the emotion data to select the scenes to extract.
[0422] Using a video editing library (FFmpeg), selected scenes are extracted from the video file and combined and edited into a highlight video.
[0423] Step 6:
[0424] The server stores the generated highlight video and generates a URL for the storage location. The URL is provided to the user, allowing the user to access the generated highlight video.
[0425] Input: highlight video
[0426] Output: Highlight video URL
[0427] Specific behavior:
[0428] Save highlight videos to a file system or cloud storage.
[0429] Generate a URL for the storage location and provide it via a notification mechanism to the user (e.g., email or in-app notification).
[0430] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0431] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0432] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0433] [Second embodiment]
[0434] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0435] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0436] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0437] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0438] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0439] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0440] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0441] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0442] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0443] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0444] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0445] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0446] The present invention is applied to a system that automatically searches for specific scenes from a video file and generates a highlight video. The system mainly includes a user, a terminal, and a server.
[0447] Basic system configuration
[0448] The system according to the present invention comprises the following main components:
[0449] 1. User's device: A device where video files are uploaded and text data specifying specific scenes is entered.
[0450] 2. Server: A computer system that uses a generative AI model to analyze video and extract scenes, then generates a highlight video.
[0451] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models to search and extract specific scenes.
[0452] Program processing flow
[0453] Upload video and enter text
[0454] Users upload video files to the server through a browser or a dedicated application. They also enter keywords and descriptions to specify specific scenes in a text input field. The device sends this data to the server as a single HTTP request.
[0455] Receiving and initial processing of input data
[0456] The server analyzes the received HTTP request and extracts the video file and text data. The video file is temporarily stored, and the text data is used for the next stage of analysis.
[0457] Launching and analyzing generative AI models
[0458] The server launches the generative AI model. First, text data is input into the natural language processing model, and keyword and semantic analysis is performed. As a specific example, the keyword "Yanagida's turn at bat" is analyzed and its features are generated.
[0459] Next, based on the generated features, the image and video analysis model scans the entire video to search for relevant scenes, and uses facial and behavioral recognition algorithms to identify the start and end times of scenes that match the analyzed features.
[0460] Scene detection and extraction
[0461] The server detects scenes based on the output data from the generative AI model. Using the time information of the identified scenes, it extracts the scenes using a video editing library (e.g., FFmpeg or OpenCV). The extracted scenes are temporarily saved as individual clips.
[0462] Highlight video generation and provision
[0463] The server combines the extracted scenes and edits them into a highlight video. The completed highlight video is saved and a URL for the saved location is generated. This URL is provided to the user, who can then view or download the generated highlight video via the link.
[0464] Specific examples
[0465] For example, if a user is searching for a scene of "Yanagida's turn at bat in a baseball broadcast in 2023," the process would be as follows: First, the user uses their device to upload a video file of the baseball broadcast to the server and enters the keyword "Yanagida's turn at bat." The server inputs the text data and video file into the generative AI model, and performs keyword analysis using a natural language processing model.
[0466] Next, an image and video analysis model scans the entire video to search for and detect scenes that correspond to Yanagi's at-bat. The identified scenes are extracted and edited to create a highlight video. Finally, users can access the generated highlight video using the provided link.
[0467] In this way, the system according to the present invention enables the user to easily search for a particular scene and realizes the process of efficiently generating a highlight video.
[0468] The processing flow will be explained below.
[0469] Step 1:
[0470] Users select video files from their own devices and upload them to the server via a browser or dedicated application.
[0471] Step 2:
[0472] The user inputs text data specifying a particular scene into the input field, for example, specifying the keyword "Yanagida's turn at bat."
[0473] Step 3:
[0474] The terminal sends the video file and text data to the server as a single HTTP request.
[0475] Step 4:
[0476] The server analyzes the received HTTP request and extracts the video file and text data. The video file is temporarily stored, and the text data is used to analyze the generative AI model.
[0477] Step 5:
[0478] The server launches the generative AI model. First, it inputs text data into the natural language processing model, analyzes the keyword "Yanagida's turn at bat," and generates its features.
[0479] Step 6:
[0480] The server inputs the generated features into an image / video analysis model, which then scans the entire video to find scenes that match the features.
[0481] Step 7:
[0482] The image and video analysis model detects the start and end times of scenes corresponding to "Yanagita's turn at bat" from the video and returns them to the server.
[0483] Step 8:
[0484] The server extracts the scene from the video file using a video editing library (e.g., FFmpeg or OpenCV) based on the time information of the detected scene.
[0485] Step 9:
[0486] The server temporarily stores the extracted scenes as one clip.
[0487] Step 10:
[0488] The server combines a plurality of extracted scenes as necessary and edits them into a highlight video.
[0489] Step 11:
[0490] The server saves the completed highlight video and generates a URL for the location where it is saved.
[0491] Step 12:
[0492] The server sends the generated URL to the terminal as an HTTP response.
[0493] Step 13:
[0494] The terminal displays the received URL to the user.
[0495] Step 14:
[0496] The user can click on the displayed link to view or download the generated highlight video.
[0497] Example 1
[0498] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0499] In modern video media, the task of efficiently extracting specific scenes from long video data and generating the highlight footage desired by users is extremely time-consuming and labor-intensive. Manually editing large amounts of video data is particularly impractical, creating a demand for automated systems. However, existing systems lack the technology for highly accurate scene detection and highlight generation. To address this issue, a system that achieves more accurate scene detection and highlight generation is needed.
[0500] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0501] In this invention, the server includes means for uploading video data, means for inputting text data specifying specific scenes, means for analyzing the video data based on the input text data using a generative AI model to detect the specific scenes, means for extracting the detected scenes and generating a highlight video, and means for providing the generated highlight video to the user. This makes it possible to automatically extract scenes specified by the user with high accuracy and easily generate a highlight video.
[0502] "Video data" refers to video files stored in digital format and has the attributes of a media file.
[0503] "Text data" is character string information that the user inputs to specify a particular scene.
[0504] A "generative AI model" is a model that uses artificial intelligence technology and includes a series of algorithms to analyze video data based on text data and detect specific scenes.
[0505] A "natural language processing model" is part of a generative AI model and is an algorithm for analyzing the meaning of input text data and extracting keywords.
[0506] The "image and video analysis model" is part of the generative AI model and is an algorithm for searching and detecting specific scenes within video data.
[0507] A "person recognition algorithm" is a technology for identifying specific people within video data.
[0508] "Movement recognition algorithms" are techniques for identifying specific movements or actions within video data.
[0509] "Cutting out" is an operation for extracting a specific portion of video data based on the start and end times of a detected scene.
[0510] A "highlight video" is a short video file created by combining specific scenes, and is a format that allows users to efficiently view scenes they want to focus on.
[0511] A "user" is an individual or group that operates the system to upload video data and specify specific scenes.
[0512] The present invention is applied to a system that automatically searches for specific scenes from video data and generates highlight videos. The system mainly includes a user, a terminal, and a server.
[0513] Basic system configuration
[0514] The system according to the present invention comprises the following main components:
[0515] 1. Terminal: A device where a user uploads video data and inputs text data specifying a particular scene.
[0516] 2. Server: A computer system that uses a generative AI model to analyze video data, extract scenes, and generate highlight footage.
[0517] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models that search for and extract specific scenes.
[0518] Example of operation
[0519] For example, if a user requests "a specific player's play from a sports broadcast in 2023," the process would proceed as follows: The user uses their device to upload a video file of the sports broadcast to the server and enters the keyword "a specific player's play." The server inputs the text data and video data into the generative AI model, and performs keyword analysis using a natural language processing model.
[0520] Next, based on the generated features, the image and video analysis model scans the entire video data to search for relevant scenes, and uses facial and behavior recognition algorithms to identify the start and end times of scenes that match the analyzed features.
[0521] Once the scenes are identified, the server uses a video editing library (e.g., FFmpeg or OpenCV) to extract the segments, which are then temporarily saved as individual clips.
[0522] Finally, the server combines the extracted scenes and edits them into a highlight video. The completed highlight video is saved on the server, and a URL for the saved location is generated. This URL is provided to users, who can then view or download the created highlight video via the link.
[0523] Hardware and software used
[0524] Hardware: User's device (PC, smartphone, etc.), server
[0525] Software: Generative AI models, video editing libraries (FFmpeg and OpenCV), natural language processing models (BERT, GPT, etc.), image and video analysis models (YOLO, etc.)
[0526] Prompt Sentence Examples
[0527] "Automatically generate highlight footage including a specific player's play from a sports broadcast in 2023."
[0528] By using this system, users can efficiently extract specific scenes from long video data and generate highlight videos in a short amount of time.
[0529] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0530] Step 1:
[0531] The user uploads video data using a device and enters text data to specify a specific scene. Specifically, the user selects video data through a browser or a dedicated application and enters keywords such as "a specific player's play." The device then sends this data to the server as a single HTTP request.
[0532] Input: Video data, text data
[0533] Output: HTTP request
[0534] Step 2:
[0535] The server analyzes the received HTTP request and extracts video and text data. The server temporarily stores the video data, and the text data is used in the next analysis stage. A parsing process is performed to analyze and extract data from the HTTP request. Keywords such as "plays by a specific player" are extracted and stored as input data.
[0536] Input: HTTP request
[0537] Output: Video data, text data
[0538] Step 3:
[0539] The server launches the generative AI model. First, the text data is input into the natural language processing model, which performs keyword analysis and semantic analysis. As a specific example, the natural language processing model analyzes the "play of a specific player" and generates its features. These features become indicators for searching the video data.
[0540] Input: Text data
[0541] Output: Text features
[0542] Step 4:
[0543] The server launches the image and video analysis model, scans the video data based on the generated text features, and searches for relevant scenes. It then uses facial and behavior recognition algorithms to identify the start and end times of specific scenes. Specifically, it analyzes the video data frame by frame to find scenes that match the features.
[0544] Input: Video data, text features
[0545] Output: Scene start and end times
[0546] Step 5:
[0547] The server uses a video editing library (e.g., FFmpeg, OpenCV) to extract the identified scene based on its time information. The extracted scene is temporarily saved as an individual clip. Specifically, the server uses a library such as FFmpeg to extract video within a specified time range.
[0548] Input: Video data, scene start and end times
[0549] Output: Video clip
[0550] Step 6:
[0551] The server combines multiple scenes and edits them into a highlight video. The completed highlight video is saved on the server and a URL for its location is generated. Specifically, an editing library such as FFmpeg is used to combine the individual clips into a single video file.
[0552] Input: Video clip
[0553] Output: highlight video file, viewing URL
[0554] Step 7:
[0555] The user can view or download the generated highlight video using the provided viewing URL. Specifically, the user opens the URL in a browser and plays or saves the video.
[0556] Input: Viewing URL
[0557] Output: Played video or downloaded video file
[0558] (Application example 1)
[0559] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0560] Conventional video search systems have had the problem that it takes a lot of time and effort to find a specific scene. Manually editing a highlight video is also cumbersome, placing a significant burden on the user. Furthermore, there has been a lack of means for efficiently viewing the generated highlight video, making it difficult to improve the user experience. The present invention aims to solve these problems by providing a system that enables users to easily search for specific scenes, generate a highlight video, and efficiently view it.
[0561] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0562] In this invention, the server includes means for uploading video files, means for inputting text data specifying specific scenes, means for analyzing the video files based on the input text data using a generative AI model and detecting specific scenes, means for extracting the detected scenes and generating a highlight video, means for using cloud storage to provide the generated highlight video to a user, and means for providing the generated highlight video by streaming, thereby enabling a user to efficiently search for specific scenes with little effort, generate a highlight video, and further watch the highlight video by streaming.
[0563] "Means for uploading video files" refers to a function that allows a user to send video data from their own device to a server and save it.
[0564] "Means for inputting text data specifying a particular scene" refers to an input interface that a user uses to specify a scene to be searched for using keywords or a description.
[0565] "Means of using a generative AI model to analyze video files based on input text data and detect specific scenes" refers to the function in which a generative AI model analyzes text data and identifies the relevant scenes in the video based on the results.
[0566] The "means for extracting detected scenes and generating a highlight video" refers to a function for cutting out identified scenes from the original video and editing them to create a highlight video.
[0567] The "means of using cloud storage to provide the generated highlight video to the user" refers to a function for storing the generated highlight video on the cloud and enabling the user to access it.
[0568] The term "means for providing the generated highlight video by streaming" refers to a function of transmitting the video via the Internet so that the user can view the highlight video generated in real time.
[0569] Basic system configuration
[0570] The system according to the present invention comprises the following main components:
[0571] 1. User's device: A device on which a video file is uploaded using a smartphone app and text data specifying a specific scene is entered.
[0572] 2. Server: A computer system that performs video analysis, scene extraction, and provides the generated highlight video to users. This server runs on Amazon AWS or other cloud platforms.
[0573] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models that search for and extract specific scenes.
[0574] Upload video and enter text
[0575] A user uploads a video file to a server through a smartphone app. Then, they enter keywords and descriptions to specify specific scenes in a text input field. For example, they enter the keyword "live performance of a popular artist." This input data is sent from the application to the server as a single HTTP request.
[0576] Receiving and initial processing of input data
[0577] The server analyzes the received HTTP request and extracts the video file and text data. The video file is temporarily stored in cloud storage such as Amazon S3, and the text data is used for the next stage of analysis.
[0578] Launching and analyzing generative AI models
[0579] The server launches a generative AI model. First, it inputs text data into a natural language processing (NLP) model (e.g., GPT-4) to perform keyword and semantic analysis. This operation identifies the content that the input keywords contain. Next, it uses an image / video analysis model (e.g., OpenCV or YOLO) to scan the entire video based on the analyzed keywords and search for relevant scenes.
[0580] Scene detection and extraction
[0581] Using the time information of the identified scene, the part is extracted using a video editing library (e.g., FFmpeg). The detected scene is temporarily saved as a clip.
[0582] Highlight video generation and provision
[0583] The server combines multiple scenes and edits them into a highlight video. The generated highlight video is saved to a server such as Amazon S3, and a URL for the saved location is generated. Users can view or download the generated highlight video from the app via this URL. The highlight video is also provided in real time via a streaming service.
[0584] Specific examples
[0585] For example, if a user searches for "live performances by popular artists," the process would proceed as follows: The user uses their device to upload a video file of the live performance to the server and enters the keyword "live performances by popular artists." The server uses a generative AI model to analyze the text data and video file, and a natural language processing model to perform keyword analysis. Next, an image and video analysis model scans the entire video to search for relevant scenes, and identified scenes are extracted as individual clips. These clips are then edited to create a highlight video, which is then provided to the user.
[0586] Prompt Sentence Examples
[0587] "Search for live performance scenes of popular artists and generate highlight videos."
[0588] In this way, the system according to the present invention realizes a process that allows a user to easily search for a particular scene and efficiently generate and view a highlight video.
[0589] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0590] Step 1:
[0591] A user uploads a video file to a server through a smartphone app. The input is a video file from the user's device, and the output is a video file stored on the server. Specifically, the user uses the app's "video upload" function to send a video file via an HTTP request. The server receives the file and stores it in cloud storage such as Amazon S3.
[0592] Step 2:
[0593] Within the same application, the user enters text data to specify a particular scene. The input is keywords and descriptions related to the particular scene, and the output is text data sent to the server. Specifically, the user enters keywords in the text field and presses the "Search" button. The text data is sent to the server along with the video file via an HTTP request.
[0594] Step 3:
[0595] The server analyzes the received video file and text data. The input is the video file and text data, and the output is the text analysis results. Specifically, the server retrieves the video file from Amazon S3 and simultaneously supplies the text data to a natural language processing (NLP) model (e.g., GPT-4). The NLP model performs keyword analysis and semantic analysis to generate the necessary features.
[0596] Step 4:
[0597] The image and video analysis part of the generative AI model is activated, and it scans the entire video based on the analyzed features to search for specific scenes. The input is the analysis results from the NLP model and the video file, and the output is the time information of the detected scene. Specifically, the generative AI model (e.g., OpenCV or YOLO) analyzes the entire video frame by frame and identifies the start and end times of the relevant scenes.
[0598] Step 5:
[0599] The server extracts the detected scenes and generates highlight clips. The input is the scene's time information and the original video file, and the output is the extracted highlight clip. Specifically, the server uses a video editing library (e.g., FFmpeg) to extract a specific scene from its start time to its end time and temporarily save it as an individual clip.
[0600] Step 6:
[0601] The server then combines the extracted scenes into a single highlight video. The input is multiple highlight clips, and the output is the completed highlight video. Specifically, the server uses FFmpeg again to combine the temporarily saved clips into a single video file.
[0602] Step 7:
[0603] The server saves the generated highlight video in cloud storage and generates a URL for the storage location. The input is the completed highlight video, and the output is the video URL. Specifically, the server uploads the completed highlight video to Amazon S3 and generates an access URL.
[0604] Step 8:
[0605] The user uses the provided URL to watch or download the generated highlight video. The input is the video URL, and the output is the highlight video played on the user's device. Specifically, the user clicks the URL link in the smartphone app and streams or downloads the video for viewing.
[0606] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0607] This invention is applied to a system that automatically searches for specific scenes from video files and generates highlight videos by combining an emotion engine that recognizes the user's emotions. The system selects scenes that correspond to the user's emotional state.
[0608] Basic system configuration
[0609] The system according to the present invention comprises the following main components:
[0610] 1. User's device: A device where video files are uploaded and text data specifying specific scenes is entered.
[0611] 2. Server: A computer system that uses a generative AI model to analyze videos and extract scenes, and combines it with the user's emotion engine to generate a highlight video.
[0612] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models to search and extract specific scenes.
[0613] 4. Emotion Engine: An algorithm to recognize the user's emotions and search and select scenes according to their emotional state.
[0614] Program processing flow
[0615] Upload video and enter text
[0616] Users upload video files to the server through a browser or a dedicated application. They also enter keywords or descriptions to identify specific scenes in a text input field. The system also collects data to recognize the user's emotions. For example, it records facial expression data and voice input while the user is watching a specific video scene.
[0617] Receiving and initial processing of input data
[0618] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, and the text data and emotion data are used for analysis by the generative AI model and emotion engine.
[0619] Launching and analyzing generative AI models
[0620] The server launches the generative AI model. First, text data is input into the natural language processing model, and keyword and semantic analysis is performed. As a specific example, the keyword "Yanagida's turn at bat" is analyzed and its features are generated.
[0621] Next, based on the generated features, the image and video analysis model scans the entire video to find relevant scenes. In addition, the emotion engine analyzes the user's emotional data and determines their emotional state.
[0622] Scene detection and extraction
[0623] The server detects scenes based on the output data from the generative AI model and the emotion analysis results from the emotion engine. For example, if the emotion engine recognizes a high level of happiness in the user, it determines that the scene is particularly important and sets a high priority for it.
[0624] Highlight video generation and provision
[0625] The server uses the time information of the detected scenes to extract those parts using a video editing library (e.g., FFmpeg or OpenCV). The extracted scenes are temporarily saved as a single clip. Multiple extracted scenes are then combined as needed to create a highlight video. The completed highlight video is saved, and a URL for its save location is generated.
[0626] Reflecting user reactions based on emotional data
[0627] The system accumulates user emotion data through the emotion engine, which can be used to improve the accuracy of future analysis results. Also, by reflecting the user's emotions toward specific scenes, the system can provide more accurate highlight videos.
[0628] Specific examples
[0629] For example, if a user is searching for a scene of "Yanagida's turn at bat in a baseball broadcast in 2023," the process would be as follows. First, the user uses their device to upload a video file of the baseball broadcast to the server and enters the keyword "Yanagida's turn at bat." The emotion engine then collects facial expressions and voice data while the user is watching the video. The server then inputs the text data and video file into a generative AI model, and performs keyword analysis using a natural language processing model.
[0630] Next, an image and video analysis model scans the entire video to search for and detect scenes that correspond to Yanagi's turn at bat. At the same time, an emotion engine analyzes the user's emotional data and prioritizes scenes in which the user expresses happiness. The identified scenes are extracted and edited to create a highlight video. Finally, the user can access the generated highlight video using the provided link.
[0631] In this way, by combining emotion engines, the present invention enables scene selection according to the individual emotional state of the user, and generates attractive content with high accuracy.
[0632] The processing flow will be explained below.
[0633] Step 1:
[0634] Users select video files from their own devices and upload them to the server via a browser or dedicated application.
[0635] Step 2:
[0636] The user inputs text data specifying a particular scene into the input field, for example, specifying the keyword "Yanagida's turn at bat."
[0637] Step 3:
[0638] The user gives permission to the system to record emotional data (facial expressions and voice) while watching the video.
[0639] Step 4:
[0640] The terminal transmits the video file, text data, and emotion data to the server as a single HTTP request.
[0641] Step 5:
[0642] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, and the text data and emotion data are used for analysis by the generative AI model and emotion engine.
[0643] Step 6:
[0644] The server launches the generative AI model. Text data is input into the natural language processing model, and semantic analysis of the text is performed. For example, the keyword "Yanagida's turn at bat" is analyzed and features are generated.
[0645] Step 7:
[0646] The server runs an image / video analysis model based on the generated features and searches the entire video for corresponding scenes. The image / video analysis model detects specific scenes using facial and behavior recognition algorithms.
[0647] Step 8:
[0648] The emotion engine analyzes the user's emotion data (facial expression data and voice data) to determine the user's emotional state, for example, identifying emotion categories such as happiness, excitement, and anxiety.
[0649] Step 9:
[0650] Based on the analysis results of the emotion engine, the server evaluates scenes in the video based on the emotion data and feature values, and sets priorities. Scenes in which the user expressed a favorable emotion are given a higher priority.
[0651] Step 10:
[0652] Based on the evaluation results, the server extracts specific scenes from the video file by identifying the start and end times of the scenes using a video editing library (such as FFmpeg or OpenCV) and extracting those parts.
[0653] Step 11:
[0654] The server temporarily stores the extracted scenes as one clip.
[0655] Step 12:
[0656] The server combines multiple extracted scenes as needed and edits them into a highlight video based on the user's emotional state.
[0657] Step 13:
[0658] The server saves the completed highlight video and generates a URL for the location where it is saved.
[0659] Step 14:
[0660] The server sends the generated URL to the terminal as an HTTP response.
[0661] Step 15:
[0662] The terminal displays the received URL to the user.
[0663] Step 16:
[0664] The user can click on the displayed link to view or download the generated highlight video.
[0665] Example 2
[0666] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0667] Conventional systems were unable to consider the user's subjective emotions when searching for specific scenes within a video file and generating a highlight video. This made it difficult to appropriately select the scenes that appealed to each individual user. Furthermore, there was a lack of technology for performing highly accurate analysis using user emotion data. This has led to a demand for improved user satisfaction.
[0668] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0669] In this invention, the server includes means for uploading video files, means for inputting text data specifying specific scenes, means for analyzing the video files based on the input text data using a generative AI model and detecting specific scenes, means for collecting user emotion data, means for analyzing the emotion data and selecting scenes based on the user's emotional state, means for extracting the detected scenes and generating a highlight video, and means for providing the generated highlight video to the user. This enables scene selection that takes user emotion into consideration and highly accurate highlight video generation.
[0670] 1. "Means for uploading video files" means the technical means that allows a user to transmit video files to a server via the Internet.
[0671] 2. "Means for inputting text data specifying a specific scene" means a technical means for providing an interface for a user to input text data such as keywords or descriptions in order to specify a specific video scene.
[0672] 3. "Generative AI model" refers to a set of algorithms that analyze input data and process it for a specific purpose, including, specifically, natural language processing models and image and video analysis models.
[0673] 4. "Natural language processing model" refers to an algorithm for analyzing text data and performing keyword analysis and semantic analysis.
[0674] 5. "Image / Video Analysis Model" refers to an algorithm that analyzes scenes in a video file and detects specific actions or objects.
[0675] 6. "Means for collecting user emotional data" refers to technical means for detecting a user's facial expressions, voice, actions, etc., and collecting emotional data based on them.
[0676] 7. "Means for analyzing emotional data and selecting scenes based on the user's emotional state" means technical means for determining the user's emotional state using collected emotional data and selecting important scenes based on that emotional state.
[0677] 8. "Means for extracting detected scenes and generating a highlight video" refers to the technical means of using a video editing library to extract that portion based on the time information of the detected scenes and edit it into a single highlight video.
[0678] 9. "Means for providing the generated highlight video to the user" refers to the technical means for saving the completed highlight video, generating a URL for the location where the video is saved, and providing the URL to the user.
[0679] In this way, by combining each of these means, it is possible to generate a highlight video that reflects the user's emotions.
[0680] The present invention relates to a system that automatically searches for specific scenes from video files and generates highlight videos by combining an emotion engine that recognizes the user's emotions. This system selects scenes that correspond to the user's emotional state.
[0681] Basic system configuration
[0682] The system according to the present invention comprises the following main components:
[0683] 1. On the user's device:
[0684] A user uses a device (such as a computer or smartphone) to upload a video file.
[0685] The device sends the video file to the server via a browser or a dedicated application.
[0686] A text entry field is used to enter keywords or descriptions to specify a particular scene.
[0687] It is also equipped with sensors to collect user emotional data, such as facial expression data and voice data.
[0688] 2. Server:
[0689] The server uses a generative AI model to analyze the video and extract scenes.
[0690] It analyzes received HTTP requests and extracts and saves video files, text data, and emotional data.
[0691] The hardware used is a computer system equipped with a high-performance processor and sufficient memory.
[0692] 3. Generative AI Model:
[0693] It consists of a set of algorithms including natural language processing models and image and video analysis models.
[0694] The natural language processing model analyzes the input text data and performs keyword analysis and semantic analysis.
[0695] The image and video analysis model uses the generated features to scan the entire video and search for relevant scenes.
[0696] 4. Emotion Engine:
[0697] An algorithm for recognizing the user's emotions and searching and selecting scenes according to their emotional state.
[0698] Facial expression and voice data are analyzed to evaluate the user's emotional state.
[0699] Example of operation
[0700] For example, if a user searches for a scene of a specific player playing in a sports broadcast in 2023, they would proceed as follows:
[0701] 1. Upload video and enter text:
[0702] A user uses his / her own terminal to upload a video file of a live sports broadcast to a server.
[0703] Enter a keyword, such as "a specific player's play," into the text entry field.
[0704] While the user is watching the video, facial expression and voice data is collected by the device's sensors.
[0705] 2. Receiving input data and initial processing:
[0706] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data.
[0707] The video file is temporarily saved, and text and emotion data are passed to the generative AI model and emotion engine.
[0708] 3. Launch and analyze the generative AI model:
[0709] The server launches the generative AI model, and the natural language processing model analyzes the text data. It performs keyword analysis based on the keyword "play by a specific player."
[0710] The image and video analysis model scans the entire video based on the generated features and searches for relevant scenes.
[0711] The emotion engine analyzes facial expression and voice data to determine the user's emotional state.
[0712] 4. Scene detection and extraction:
[0713] The server detects scenes based on the output data from the generative AI model and the emotion analysis results from the emotion engine.
[0714] If the emotion engine rates the user's happiness highly, it sets the priority of that scene high.
[0715] 5. Highlight video generation and delivery:
[0716] The server extracts the detected scene using a video editing library (e.g., FFmpeg or OpenCV) based on the time information of the detected scene.
[0717] Multiple extracted scenes are combined and edited into a highlight video.
[0718] The completed highlight video is saved, and a URL for the saved location is generated and provided to the user.
[0719] In this way, by combining emotion engines, the present invention makes it possible to select scenes according to the individual emotional state of the user, and to generate attractive content with high accuracy.
[0720] Prompt Sentence Examples
[0721] It searches for and extracts specific scenes based on text data entered by the user, such as "a specific player's play from a sports broadcast in 2023."
[0722] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0723] Step 1: Upload video and enter text
[0724] Users select video files from a dedicated application or browser on their own devices and upload them to the server. Users enter keywords or descriptions in a text input field to specify specific scenes. Emotional data such as facial expressions and voice data are also collected while the user is watching the video.
[0725] Input: Video file, text data (e.g., "A specific player's play"), emotional data (facial expressions, voice)
[0726] Output: The video file, text data, and emotion data sent by the user are uploaded to the server.
[0727] Specific operation: The user selects a video file and clicks the upload button. Next, the user enters a keyword specifying a specific scene in the text input field and clicks the send button. While watching the video, the user's facial expression data and voice data are collected by the device's sensors.
[0728] Step 2: Receiving input data and initial processing
[0729] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, and the text data and emotion data are prepared for analysis by the generative AI model and emotion engine.
[0730] Input: Video files uploaded to the server, text data, emotion data
[0731] Output: Temporarily saved video files, analyzable text data and emotion data
[0732] Specific operation: The server analyzes the received HTTP request and saves the video file in the specified directory. Next, it stores the text data and emotion data in memory and prepares for analysis.
[0733] Step 3: Launch and analyze the generative AI model
[0734] The server launches the generative AI model. Text data is input into the natural language processing model, which performs keyword and semantic analysis. Based on the analyzed features, the image and video analysis model scans the entire video to search for relevant scenes. At the same time, the emotion engine analyzes the user's emotional data and determines their emotional state.
[0735] Input: text data, video files, emotion data
[0736] Output: Keyword analysis results, candidate scenes, and user's emotional state
[0737] How it works: The server analyzes the text data entered into the natural language processing model and analyzes the keyword "a specific player's play." Next, based on the generated features, the image and video analysis model scans the entire video to search for specific scenes. At the same time, the emotion engine analyzes facial expression and voice data to evaluate the user's emotions.
[0738] Step 4: Scene detection and extraction
[0739] The server detects scenes based on the output data of the generative AI model and the emotion analysis results of the emotion engine. The emotion engine selects scenes that evoke a specific emotion (e.g., happiness) and sets a priority.
[0740] Input: Analysis results of the generative AI model, analysis results of the emotion engine
[0741] Output: Time information of detected scenes, prioritized scenes
[0742] Specific operation: The server identifies the start and end times of scenes corresponding to Yanagida's at-bats, sets priorities based on the results of emotion analysis, and adds high-priority scenes to a list.
[0743] Step 5: Generate and serve highlight videos
[0744] The server uses a video editing library (e.g., FFmpeg or OpenCV) to extract the relevant parts based on the time information of the detected scenes. If necessary, multiple scenes are combined and edited into a highlight video. The completed highlight video is saved, and a URL for the saved location is generated and provided to the user.
[0745] Input: Time information of detected scenes, prioritized scenes
[0746] Output: Saved highlight video, access URL to the highlight video
[0747] Specific operation: The server uses FFmpeg commands to extract specific scenes from a video file, combines multiple scenes into a single highlight video, and saves the generated highlight video to storage. Finally, it generates a URL for the highlight video and notifies the user.
[0748] (Application example 2)
[0749] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0750] In conventional highlight video generation systems, the detection and extraction of specific scenes is often performed solely through analysis of text data, without taking into account the user's emotions or preferences, resulting in the generated highlight videos not necessarily meeting the user's expectations. Furthermore, personalized content that reflects the user's emotions is in demand, particularly in entertainment applications, but few systems have this functionality.
[0751] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0752] In this invention, the server includes means for uploading video files, means for inputting text data specifying specific scenes, means for analyzing the video files based on the input text data using a generative AI model to detect specific scenes, means for extracting the detected scenes and generating a highlight video, means for collecting user emotion data, means for determining the importance of scenes based on the emotion data, and means for providing the generated highlight video to the user, thereby enabling the generation of a personalized highlight video that takes user emotions into consideration.
[0753] A "video file" is a digital file containing visual and audio information.
[0754] "Uploading" is the act of transferring data from a user's terminal to a remote computer such as a server.
[0755] "Text data" is digital data that contains written information.
[0756] A "generative AI model" is a set of artificial intelligence algorithms designed to perform natural language processing and image and video analysis.
[0757] An "emotion engine" is an algorithm that analyzes the user's emotional state and performs processing based on that data.
[0758] A "highlight video" is a shortened version of a video that has been edited to extract specific scenes.
[0759] "User" means a person or organization that uses the system.
[0760] A "natural language processing model" is an artificial intelligence model for understanding and analyzing text data.
[0761] An "image and video analysis model" is an artificial intelligence model for analyzing and recognizing the content of images and videos.
[0762] "Facial recognition" is a technology that detects and identifies human faces from images and videos.
[0763] "Behavior recognition" is a technology that detects and identifies human actions and movements from images and videos.
[0764] "Emotion data" is digital data that expresses the user's emotional state using numerical values and categories.
[0765] "Decision making" is the process of reaching a conclusion based on data.
[0766] System Overview
[0767] This invention relates to a system that recognizes a user's emotions, automatically searches and extracts specific scenes from videos based on those emotions, and generates highlight videos. The system consists of the following main components:
[0768] 1. Terminal: A device where a user uploads video files and enters text data specifying specific scenes. Typically, this is a smartphone or tablet.
[0769] 2. Server: A computer system that uses a generative AI model to analyze video files, detect and extract specific scenes, and generate a highlight video.
[0770] 3. Generative AI model: A set of algorithms, including natural language processing models and image and video analysis models, used to search and extract specific scenes.
[0771] 4. Emotion engine: An algorithm that analyzes the user's emotional data and determines the importance of a scene based on that data.
[0772] System Operation
[0773] Upload video and enter keywords
[0774] Users upload video files to the server via a smartphone app. They input keywords and descriptions in text format to specify specific scenes. While the user is watching the video, the camera and microphone are used to collect facial expressions and voice data.
[0775] Data reception and initial processing
[0776] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily saved, and the text data is input into a natural language processing model. The emotion data is sent to the emotion engine.
[0777] Analysis of generative AI models
[0778] The server launches a generative AI model to analyze the text data. A natural language processing model (e.g., TensorFlow model) performs keyword and semantic analysis and generates features. Next, an image and video analysis model scans the entire video based on the generated features to search for relevant scenes. At the same time, an emotion engine analyzes the user's emotional data and determines their emotional state (e.g., Microsoft Azure Emotion API).
[0779] Scene detection and extraction
[0780] Based on the analysis results of the generative AI model and emotion engine, the server detects specific scenes. It determines the importance of each scene based on the emotion data and extracts high-priority scenes. It then uses a video editing library (e.g., FFmpeg) to extract the detected scenes and generate a highlight video.
[0781] Highlight video provided
[0782] The generated highlight video is stored on the server and a generated link is provided to the user, who can use the link to view and share the highlight video.
[0783] Specific examples
[0784] For example, if a user wants to search for "scenes of a specific player from a sports broadcast in 2023," the process would proceed as follows: The user uploads a video file of the sports broadcast to the server using their device and enters the keyword "scenes of a specific player." The emotion engine also collects facial expressions and voice data while the user is watching the video. The server inputs the text data and video file into the generative AI model, and performs keyword analysis using a natural language processing model. The image and video analysis model then scans the entire video to search for and detect relevant scenes. At the same time, the emotion engine analyzes the user's emotional data and selects scenes with high importance. The identified scenes are then extracted and edited to create a highlight video. Finally, the user can access the highlight video using the generated link.
[0785] Examples of prompt statements
[0786] "Look for scenes that include a specific player."
[0787] "Create important highlights based on the emotional data of this scene."
[0788] This system enables the generation of personalized highlight videos that reflect the user's emotions.
[0789] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0790] Step 1:
[0791] A user uploads a video file using a smartphone app and inputs text data (keywords and descriptions) to specify a specific scene. While the user is watching the video, the device uses a camera and microphone to collect facial expressions and voice data. The collected data is then sent to a server.
[0792] Input: Video files, text data, facial expressions and audio data
[0793] Output: Send data to the server
[0794] Specific behavior:
[0795] The application accepts user operations and displays a UI that allows the user to select a video file.
[0796] Provide a text field for entering keywords and descriptions.
[0797] It asks for permission to use the camera and microphone, and if the user agrees, it begins collecting data.
[0798] Step 2:
[0799] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, the text data is sent to a natural language processing model (e.g., TensorFlow model), and the emotion data is sent to an emotion engine (e.g., Microsoft Azure Emotion API).
[0800] Input: HTTP request
[0801] Output: Video file, text data, emotion data extraction
[0802] Specific behavior:
[0803] The server receives the HTTP request and parses the video file and text data.
[0804] Temporarily save video files to a file system or cloud storage.
[0805] The text data and emotion data are sent to the corresponding analysis units.
[0806] Step 3:
[0807] The generative AI model analyzes the text data. The natural language processing model performs keyword and semantic analysis and generates feature vectors. Based on these feature vectors, the image and video analysis model scans the entire video and searches for relevant scenes.
[0808] Input: Text data
[0809] Output: Features, list of detected scenes
[0810] Specific behavior:
[0811] A natural language processing model receives the text data and analyzes it for keywords and meaning.
[0812] Features are generated and the data is passed to an image / video analysis model.
[0813] The model scans the video file and lists the start and end times of the relevant scenes.
[0814] Step 4:
[0815] The emotion engine analyzes the user's emotion data and determines the emotional state the user is in. Based on the analysis results, the importance of the scene is determined.
[0816] Input: Emotion data
[0817] Output: Emotional state analysis result, scene importance judgment
[0818] Specific behavior:
[0819] The emotion engine analyzes the user's facial expressions and voice data to detect emotional states such as happiness, surprise, and excitement.
[0820] Emotional data is collected for each scene and the importance of that scene is calculated.
[0821] Step 5:
[0822] The server detects and extracts specific scenes based on the analysis results of the generated AI model and emotion engine, and uses a video editing library (e.g., FFmpeg) to extract the detected scenes and generate a highlight video.
[0823] Input: List of scenes, emotional state analysis results
[0824] Output: Highlight video
[0825] Specific behavior:
[0826] The list of detected scenes is compared with the emotion data to select the scenes to extract.
[0827] Using a video editing library (FFmpeg), selected scenes are extracted from the video file and combined and edited into a highlight video.
[0828] Step 6:
[0829] The server stores the generated highlight video and generates a URL for the storage location. The URL is provided to the user, allowing the user to access the generated highlight video.
[0830] Input: highlight video
[0831] Output: Highlight video URL
[0832] Specific behavior:
[0833] Save highlight videos to a file system or cloud storage.
[0834] Generate a URL for the storage location and provide it via a notification mechanism to the user (e.g., email or in-app notification).
[0835] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0836] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0837] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0838] [Third embodiment]
[0839] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0840] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0841] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0842] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0843] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0844] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0845] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0846] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0847] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0848] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0849] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0850] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0851] The present invention is applied to a system that automatically searches for specific scenes from a video file and generates a highlight video. The system mainly includes a user, a terminal, and a server.
[0852] Basic system configuration
[0853] The system according to the present invention comprises the following main components:
[0854] 1. User's device: A device where video files are uploaded and text data specifying specific scenes is entered.
[0855] 2. Server: A computer system that uses a generative AI model to analyze video and extract scenes, then generates a highlight video.
[0856] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models to search and extract specific scenes.
[0857] Program processing flow
[0858] Upload video and enter text
[0859] Users upload video files to the server through a browser or a dedicated application. They also enter keywords and descriptions to specify specific scenes in a text input field. The device sends this data to the server as a single HTTP request.
[0860] Receiving and initial processing of input data
[0861] The server analyzes the received HTTP request and extracts the video file and text data. The video file is temporarily stored, and the text data is used for the next stage of analysis.
[0862] Launching and analyzing generative AI models
[0863] The server launches the generative AI model. First, text data is input into the natural language processing model, and keyword and semantic analysis is performed. As a specific example, the keyword "Yanagida's turn at bat" is analyzed and its features are generated.
[0864] Next, based on the generated features, the image and video analysis model scans the entire video to search for relevant scenes, and uses facial and behavioral recognition algorithms to identify the start and end times of scenes that match the analyzed features.
[0865] Scene detection and extraction
[0866] The server detects scenes based on the output data from the generative AI model. Using the time information of the identified scenes, it extracts the scenes using a video editing library (e.g., FFmpeg or OpenCV). The extracted scenes are temporarily saved as individual clips.
[0867] Highlight video generation and provision
[0868] The server combines the extracted scenes and edits them into a highlight video. The completed highlight video is saved and a URL for the saved location is generated. This URL is provided to the user, who can then view or download the generated highlight video via the link.
[0869] Specific examples
[0870] For example, if a user is searching for a scene of "Yanagida's turn at bat in a baseball broadcast in 2023," the process would be as follows: First, the user uses their device to upload a video file of the baseball broadcast to the server and enters the keyword "Yanagida's turn at bat." The server inputs the text data and video file into the generative AI model, and performs keyword analysis using a natural language processing model.
[0871] Next, an image and video analysis model scans the entire video to search for and detect scenes that correspond to Yanagi's at-bat. The identified scenes are extracted and edited to create a highlight video. Finally, users can access the generated highlight video using the provided link.
[0872] In this way, the system according to the present invention enables the user to easily search for a particular scene and realizes the process of efficiently generating a highlight video.
[0873] The processing flow will be explained below.
[0874] Step 1:
[0875] Users select video files from their own devices and upload them to the server via a browser or dedicated application.
[0876] Step 2:
[0877] The user inputs text data specifying a particular scene into the input field, for example, specifying the keyword "Yanagida's turn at bat."
[0878] Step 3:
[0879] The terminal sends the video file and text data to the server as a single HTTP request.
[0880] Step 4:
[0881] The server analyzes the received HTTP request and extracts the video file and text data. The video file is temporarily stored, and the text data is used to analyze the generative AI model.
[0882] Step 5:
[0883] The server launches the generative AI model. First, it inputs text data into the natural language processing model, analyzes the keyword "Yanagida's turn at bat," and generates its features.
[0884] Step 6:
[0885] The server inputs the generated features into an image / video analysis model, which then scans the entire video to find scenes that match the features.
[0886] Step 7:
[0887] The image and video analysis model detects the start and end times of scenes corresponding to "Yanagita's turn at bat" from the video and returns them to the server.
[0888] Step 8:
[0889] The server extracts the scene from the video file using a video editing library (e.g., FFmpeg or OpenCV) based on the time information of the detected scene.
[0890] Step 9:
[0891] The server temporarily stores the extracted scenes as one clip.
[0892] Step 10:
[0893] The server combines a plurality of extracted scenes as necessary and edits them into a highlight video.
[0894] Step 11:
[0895] The server saves the completed highlight video and generates a URL for the location where it is saved.
[0896] Step 12:
[0897] The server sends the generated URL to the terminal as an HTTP response.
[0898] Step 13:
[0899] The terminal displays the received URL to the user.
[0900] Step 14:
[0901] The user can click on the displayed link to view or download the generated highlight video.
[0902] Example 1
[0903] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0904] In modern video media, the task of efficiently extracting specific scenes from long video data and generating the highlight footage desired by users is extremely time-consuming and labor-intensive. Manually editing large amounts of video data is particularly impractical, creating a demand for automated systems. However, existing systems lack the technology for highly accurate scene detection and highlight generation. To address this issue, a system that achieves more accurate scene detection and highlight generation is needed.
[0905] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0906] In this invention, the server includes means for uploading video data, means for inputting text data specifying specific scenes, means for analyzing the video data based on the input text data using a generative AI model to detect the specific scenes, means for extracting the detected scenes and generating a highlight video, and means for providing the generated highlight video to the user. This makes it possible to automatically extract scenes specified by the user with high accuracy and easily generate a highlight video.
[0907] "Video data" refers to video files stored in digital format and has the attributes of a media file.
[0908] "Text data" is character string information that the user inputs to specify a particular scene.
[0909] A "generative AI model" is a model that uses artificial intelligence technology and includes a series of algorithms to analyze video data based on text data and detect specific scenes.
[0910] A "natural language processing model" is part of a generative AI model and is an algorithm for analyzing the meaning of input text data and extracting keywords.
[0911] The "image and video analysis model" is part of the generative AI model and is an algorithm for searching and detecting specific scenes within video data.
[0912] A "person recognition algorithm" is a technology for identifying specific people within video data.
[0913] "Movement recognition algorithms" are techniques for identifying specific movements or actions within video data.
[0914] "Cutting out" is an operation for extracting a specific portion of video data based on the start and end times of a detected scene.
[0915] A "highlight video" is a short video file created by combining specific scenes, and is a format that allows users to efficiently view scenes they want to focus on.
[0916] A "user" is an individual or group that operates the system to upload video data and specify specific scenes.
[0917] The present invention is applied to a system that automatically searches for specific scenes from video data and generates highlight videos. The system mainly includes a user, a terminal, and a server.
[0918] Basic system configuration
[0919] The system according to the present invention comprises the following main components:
[0920] 1. Terminal: A device where a user uploads video data and inputs text data specifying a particular scene.
[0921] 2. Server: A computer system that uses a generative AI model to analyze video data, extract scenes, and generate highlight footage.
[0922] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models that search for and extract specific scenes.
[0923] Example of operation
[0924] For example, if a user requests "a specific player's play from a sports broadcast in 2023," the process would proceed as follows: The user uses their device to upload a video file of the sports broadcast to the server and enters the keyword "a specific player's play." The server inputs the text data and video data into the generative AI model, and performs keyword analysis using a natural language processing model.
[0925] Next, based on the generated features, the image and video analysis model scans the entire video data to search for relevant scenes, and uses facial and behavior recognition algorithms to identify the start and end times of scenes that match the analyzed features.
[0926] Once the scenes are identified, the server uses a video editing library (e.g., FFmpeg or OpenCV) to extract the segments, which are then temporarily saved as individual clips.
[0927] Finally, the server combines the extracted scenes and edits them into a highlight video. The completed highlight video is saved on the server, and a URL for the saved location is generated. This URL is provided to users, who can then view or download the created highlight video via the link.
[0928] Hardware and software used
[0929] Hardware: User's device (PC, smartphone, etc.), server
[0930] Software: Generative AI models, video editing libraries (FFmpeg and OpenCV), natural language processing models (BERT, GPT, etc.), image and video analysis models (YOLO, etc.)
[0931] Prompt Sentence Examples
[0932] "Automatically generate highlight footage including a specific player's play from a sports broadcast in 2023."
[0933] By using this system, users can efficiently extract specific scenes from long video data and generate highlight videos in a short amount of time.
[0934] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0935] Step 1:
[0936] The user uploads video data using a device and enters text data to specify a specific scene. Specifically, the user selects video data through a browser or a dedicated application and enters keywords such as "a specific player's play." The device then sends this data to the server as a single HTTP request.
[0937] Input: Video data, text data
[0938] Output: HTTP request
[0939] Step 2:
[0940] The server analyzes the received HTTP request and extracts video and text data. The server temporarily stores the video data, and the text data is used in the next analysis stage. A parsing process is performed to analyze and extract data from the HTTP request. Keywords such as "plays by a specific player" are extracted and stored as input data.
[0941] Input: HTTP request
[0942] Output: Video data, text data
[0943] Step 3:
[0944] The server launches the generative AI model. First, the text data is input into the natural language processing model, which performs keyword analysis and semantic analysis. As a specific example, the natural language processing model analyzes the "play of a specific player" and generates its features. These features become indicators for searching the video data.
[0945] Input: Text data
[0946] Output: Text features
[0947] Step 4:
[0948] The server launches the image and video analysis model, scans the video data based on the generated text features, and searches for relevant scenes. It then uses facial and behavior recognition algorithms to identify the start and end times of specific scenes. Specifically, it analyzes the video data frame by frame to find scenes that match the features.
[0949] Input: Video data, text features
[0950] Output: Scene start and end times
[0951] Step 5:
[0952] The server uses a video editing library (e.g., FFmpeg, OpenCV) to extract the identified scene based on its time information. The extracted scene is temporarily saved as an individual clip. Specifically, the server uses a library such as FFmpeg to extract video within a specified time range.
[0953] Input: Video data, scene start and end times
[0954] Output: Video clip
[0955] Step 6:
[0956] The server combines multiple scenes and edits them into a highlight video. The completed highlight video is saved on the server and a URL for its location is generated. Specifically, an editing library such as FFmpeg is used to combine the individual clips into a single video file.
[0957] Input: Video clip
[0958] Output: highlight video file, viewing URL
[0959] Step 7:
[0960] The user can view or download the generated highlight video using the provided viewing URL. Specifically, the user opens the URL in a browser and plays or saves the video.
[0961] Input: Viewing URL
[0962] Output: Played video or downloaded video file
[0963] (Application example 1)
[0964] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0965] Conventional video search systems have had the problem that it takes a lot of time and effort to find a specific scene. Manually editing a highlight video is also cumbersome, placing a significant burden on the user. Furthermore, there has been a lack of means for efficiently viewing the generated highlight video, making it difficult to improve the user experience. The present invention aims to solve these problems by providing a system that enables users to easily search for specific scenes, generate a highlight video, and efficiently view it.
[0966] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0967] In this invention, the server includes means for uploading video files, means for inputting text data specifying specific scenes, means for analyzing the video files based on the input text data using a generative AI model and detecting specific scenes, means for extracting the detected scenes and generating a highlight video, means for using cloud storage to provide the generated highlight video to a user, and means for providing the generated highlight video by streaming, thereby enabling a user to efficiently search for specific scenes with little effort, generate a highlight video, and further watch the highlight video by streaming.
[0968] "Means for uploading video files" refers to a function that allows a user to send video data from their own device to a server and save it.
[0969] "Means for inputting text data specifying a particular scene" refers to an input interface that a user uses to specify a scene to be searched for using keywords or a description.
[0970] "Means of using a generative AI model to analyze video files based on input text data and detect specific scenes" refers to the function in which a generative AI model analyzes text data and identifies the relevant scenes in the video based on the results.
[0971] The "means for extracting detected scenes and generating a highlight video" refers to a function for cutting out identified scenes from the original video and editing them to create a highlight video.
[0972] The "means of using cloud storage to provide the generated highlight video to the user" refers to a function for storing the generated highlight video on the cloud and enabling the user to access it.
[0973] The term "means for providing the generated highlight video by streaming" refers to a function of transmitting the video via the Internet so that the user can view the highlight video generated in real time.
[0974] Basic system configuration
[0975] The system according to the present invention comprises the following main components:
[0976] 1. User's device: A device on which a video file is uploaded using a smartphone app and text data specifying a specific scene is entered.
[0977] 2. Server: A computer system that performs video analysis, scene extraction, and provides the generated highlight video to users. This server runs on Amazon AWS or other cloud platforms.
[0978] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models that search for and extract specific scenes.
[0979] Upload video and enter text
[0980] A user uploads a video file to a server through a smartphone app. Then, they enter keywords and descriptions to specify specific scenes in a text input field. For example, they enter the keyword "live performance of a popular artist." This input data is sent from the application to the server as a single HTTP request.
[0981] Receiving and initial processing of input data
[0982] The server analyzes the received HTTP request and extracts the video file and text data. The video file is temporarily stored in cloud storage such as Amazon S3, and the text data is used for the next stage of analysis.
[0983] Launching and analyzing generative AI models
[0984] The server launches a generative AI model. First, it inputs text data into a natural language processing (NLP) model (e.g., GPT-4) to perform keyword and semantic analysis. This operation identifies the content that the input keywords contain. Next, it uses an image / video analysis model (e.g., OpenCV or YOLO) to scan the entire video based on the analyzed keywords and search for relevant scenes.
[0985] Scene detection and extraction
[0986] Using the time information of the identified scene, the part is extracted using a video editing library (e.g., FFmpeg). The detected scene is temporarily saved as a clip.
[0987] Highlight video generation and provision
[0988] The server combines multiple scenes and edits them into a highlight video. The generated highlight video is saved to a server such as Amazon S3, and a URL for the saved location is generated. Users can view or download the generated highlight video from the app via this URL. The highlight video is also provided in real time via a streaming service.
[0989] Specific examples
[0990] For example, if a user searches for "live performances by popular artists," the process would proceed as follows: The user uses their device to upload a video file of the live performance to the server and enters the keyword "live performances by popular artists." The server uses a generative AI model to analyze the text data and video file, and a natural language processing model to perform keyword analysis. Next, an image and video analysis model scans the entire video to search for relevant scenes, and identified scenes are extracted as individual clips. These clips are then edited to create a highlight video, which is then provided to the user.
[0991] Prompt Sentence Examples
[0992] "Search for live performance scenes of popular artists and generate highlight videos."
[0993] In this way, the system according to the present invention realizes a process that allows a user to easily search for a particular scene and efficiently generate and view a highlight video.
[0994] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0995] Step 1:
[0996] A user uploads a video file to a server through a smartphone app. The input is a video file from the user's device, and the output is a video file stored on the server. Specifically, the user uses the app's "video upload" function to send a video file via an HTTP request. The server receives the file and stores it in cloud storage such as Amazon S3.
[0997] Step 2:
[0998] Within the same application, the user enters text data to specify a particular scene. The input is keywords and descriptions related to the particular scene, and the output is text data sent to the server. Specifically, the user enters keywords in the text field and presses the "Search" button. The text data is sent to the server along with the video file via an HTTP request.
[0999] Step 3:
[1000] The server analyzes the received video file and text data. The input is the video file and text data, and the output is the text analysis results. Specifically, the server retrieves the video file from Amazon S3 and simultaneously supplies the text data to a natural language processing (NLP) model (e.g., GPT-4). The NLP model performs keyword analysis and semantic analysis to generate the necessary features.
[1001] Step 4:
[1002] The image and video analysis part of the generative AI model is activated, and it scans the entire video based on the analyzed features to search for specific scenes. The input is the analysis results from the NLP model and the video file, and the output is the time information of the detected scene. Specifically, the generative AI model (e.g., OpenCV or YOLO) analyzes the entire video frame by frame and identifies the start and end times of the relevant scenes.
[1003] Step 5:
[1004] The server extracts the detected scenes and generates highlight clips. The input is the scene's time information and the original video file, and the output is the extracted highlight clip. Specifically, the server uses a video editing library (e.g., FFmpeg) to extract a specific scene from its start time to its end time and temporarily save it as an individual clip.
[1005] Step 6:
[1006] The server then combines the extracted scenes into a single highlight video. The input is multiple highlight clips, and the output is the completed highlight video. Specifically, the server uses FFmpeg again to combine the temporarily saved clips into a single video file.
[1007] Step 7:
[1008] The server saves the generated highlight video in cloud storage and generates a URL for the storage location. The input is the completed highlight video, and the output is the video URL. Specifically, the server uploads the completed highlight video to Amazon S3 and generates an access URL.
[1009] Step 8:
[1010] The user uses the provided URL to watch or download the generated highlight video. The input is the video URL, and the output is the highlight video played on the user's device. Specifically, the user clicks the URL link in the smartphone app and streams or downloads the video for viewing.
[1011] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1012] This invention is applied to a system that automatically searches for specific scenes from video files and generates highlight videos by combining an emotion engine that recognizes the user's emotions. The system selects scenes that correspond to the user's emotional state.
[1013] Basic system configuration
[1014] The system according to the present invention comprises the following main components:
[1015] 1. User's device: A device where video files are uploaded and text data specifying specific scenes is entered.
[1016] 2. Server: A computer system that uses a generative AI model to analyze videos and extract scenes, and combines it with the user's emotion engine to generate a highlight video.
[1017] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models to search and extract specific scenes.
[1018] 4. Emotion Engine: An algorithm to recognize the user's emotions and search and select scenes according to their emotional state.
[1019] Program processing flow
[1020] Upload video and enter text
[1021] Users upload video files to the server through a browser or a dedicated application. They also enter keywords or descriptions to identify specific scenes in a text input field. The system also collects data to recognize the user's emotions. For example, it records facial expression data and voice input while the user is watching a specific video scene.
[1022] Receiving and initial processing of input data
[1023] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, and the text data and emotion data are used for analysis by the generative AI model and emotion engine.
[1024] Launching and analyzing generative AI models
[1025] The server launches the generative AI model. First, text data is input into the natural language processing model, and keyword and semantic analysis is performed. As a specific example, the keyword "Yanagida's turn at bat" is analyzed and its features are generated.
[1026] Next, based on the generated features, the image and video analysis model scans the entire video to find relevant scenes. In addition, the emotion engine analyzes the user's emotional data and determines their emotional state.
[1027] Scene detection and extraction
[1028] The server detects scenes based on the output data from the generative AI model and the emotion analysis results from the emotion engine. For example, if the emotion engine recognizes a high level of happiness in the user, it determines that the scene is particularly important and sets a high priority for it.
[1029] Highlight video generation and provision
[1030] The server uses the time information of the detected scenes to extract those parts using a video editing library (e.g., FFmpeg or OpenCV). The extracted scenes are temporarily saved as a single clip. Multiple extracted scenes are then combined as needed to create a highlight video. The completed highlight video is saved, and a URL for its save location is generated.
[1031] Reflecting user reactions based on emotional data
[1032] The system accumulates user emotion data through the emotion engine, which can be used to improve the accuracy of future analysis results. Also, by reflecting the user's emotions toward specific scenes, the system can provide more accurate highlight videos.
[1033] Specific examples
[1034] For example, if a user is searching for a scene of "Yanagida's turn at bat in a baseball broadcast in 2023," the process would be as follows. First, the user uses their device to upload a video file of the baseball broadcast to the server and enters the keyword "Yanagida's turn at bat." The emotion engine then collects facial expressions and voice data while the user is watching the video. The server then inputs the text data and video file into a generative AI model, and performs keyword analysis using a natural language processing model.
[1035] Next, an image and video analysis model scans the entire video to search for and detect scenes that correspond to Yanagi's turn at bat. At the same time, an emotion engine analyzes the user's emotional data and prioritizes scenes in which the user expresses happiness. The identified scenes are extracted and edited to create a highlight video. Finally, the user can access the generated highlight video using the provided link.
[1036] In this way, by combining emotion engines, the present invention enables scene selection according to the individual emotional state of the user, and generates attractive content with high accuracy.
[1037] The processing flow will be explained below.
[1038] Step 1:
[1039] Users select video files from their own devices and upload them to the server via a browser or dedicated application.
[1040] Step 2:
[1041] The user inputs text data specifying a particular scene into the input field, for example, specifying the keyword "Yanagida's turn at bat."
[1042] Step 3:
[1043] The user gives permission to the system to record emotional data (facial expressions and voice) while watching the video.
[1044] Step 4:
[1045] The terminal transmits the video file, text data, and emotion data to the server as a single HTTP request.
[1046] Step 5:
[1047] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, and the text data and emotion data are used for analysis by the generative AI model and emotion engine.
[1048] Step 6:
[1049] The server launches the generative AI model. Text data is input into the natural language processing model, and semantic analysis of the text is performed. For example, the keyword "Yanagida's turn at bat" is analyzed and features are generated.
[1050] Step 7:
[1051] The server runs an image / video analysis model based on the generated features and searches the entire video for corresponding scenes. The image / video analysis model detects specific scenes using facial and behavior recognition algorithms.
[1052] Step 8:
[1053] The emotion engine analyzes the user's emotion data (facial expression data and voice data) to determine the user's emotional state, for example, identifying emotion categories such as happiness, excitement, and anxiety.
[1054] Step 9:
[1055] Based on the analysis results of the emotion engine, the server evaluates scenes in the video based on the emotion data and feature values, and sets priorities. Scenes in which the user expressed a favorable emotion are given a higher priority.
[1056] Step 10:
[1057] Based on the evaluation results, the server extracts specific scenes from the video file by identifying the start and end times of the scenes using a video editing library (such as FFmpeg or OpenCV) and extracting those parts.
[1058] Step 11:
[1059] The server temporarily stores the extracted scenes as one clip.
[1060] Step 12:
[1061] The server combines multiple extracted scenes as needed and edits them into a highlight video based on the user's emotional state.
[1062] Step 13:
[1063] The server saves the completed highlight video and generates a URL for the location where it is saved.
[1064] Step 14:
[1065] The server sends the generated URL to the terminal as an HTTP response.
[1066] Step 15:
[1067] The terminal displays the received URL to the user.
[1068] Step 16:
[1069] The user can click on the displayed link to view or download the generated highlight video.
[1070] Example 2
[1071] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1072] Conventional systems were unable to consider the user's subjective emotions when searching for specific scenes within a video file and generating a highlight video. This made it difficult to appropriately select the scenes that appealed to each individual user. Furthermore, there was a lack of technology for performing highly accurate analysis using user emotion data. This has led to a demand for improved user satisfaction.
[1073] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1074] In this invention, the server includes means for uploading video files, means for inputting text data specifying specific scenes, means for analyzing the video files based on the input text data using a generative AI model and detecting specific scenes, means for collecting user emotion data, means for analyzing the emotion data and selecting scenes based on the user's emotional state, means for extracting the detected scenes and generating a highlight video, and means for providing the generated highlight video to the user. This enables scene selection that takes user emotion into consideration and highly accurate highlight video generation.
[1075] 1. "Means for uploading video files" means the technical means that allows a user to transmit video files to a server via the Internet.
[1076] 2. "Means for inputting text data specifying a specific scene" means a technical means for providing an interface for a user to input text data such as keywords or descriptions in order to specify a specific video scene.
[1077] 3. "Generative AI model" refers to a set of algorithms that analyze input data and process it for a specific purpose, including, specifically, natural language processing models and image and video analysis models.
[1078] 4. "Natural language processing model" refers to an algorithm for analyzing text data and performing keyword analysis and semantic analysis.
[1079] 5. "Image / Video Analysis Model" refers to an algorithm that analyzes scenes in a video file and detects specific actions or objects.
[1080] 6. "Means for collecting user emotional data" refers to technical means for detecting a user's facial expressions, voice, actions, etc., and collecting emotional data based on them.
[1081] 7. "Means for analyzing emotional data and selecting scenes based on the user's emotional state" means technical means for determining the user's emotional state using collected emotional data and selecting important scenes based on that emotional state.
[1082] 8. "Means for extracting detected scenes and generating a highlight video" refers to the technical means of using a video editing library to extract that portion based on the time information of the detected scenes and edit it into a single highlight video.
[1083] 9. "Means for providing the generated highlight video to the user" refers to the technical means for saving the completed highlight video, generating a URL for the location where the video is saved, and providing the URL to the user.
[1084] In this way, by combining each of these means, it is possible to generate a highlight video that reflects the user's emotions.
[1085] The present invention relates to a system that automatically searches for specific scenes from video files and generates highlight videos by combining an emotion engine that recognizes the user's emotions. This system selects scenes that correspond to the user's emotional state.
[1086] Basic system configuration
[1087] The system according to the present invention comprises the following main components:
[1088] 1. On the user's device:
[1089] A user uses a device (such as a computer or smartphone) to upload a video file.
[1090] The device sends the video file to the server via a browser or a dedicated application.
[1091] A text entry field is used to enter keywords or descriptions to specify a particular scene.
[1092] It is also equipped with sensors to collect user emotional data, such as facial expression data and voice data.
[1093] 2. Server:
[1094] The server uses a generative AI model to analyze the video and extract scenes.
[1095] It analyzes received HTTP requests and extracts and saves video files, text data, and emotional data.
[1096] The hardware used is a computer system equipped with a high-performance processor and sufficient memory.
[1097] 3. Generative AI Model:
[1098] It consists of a set of algorithms including natural language processing models and image and video analysis models.
[1099] The natural language processing model analyzes the input text data and performs keyword analysis and semantic analysis.
[1100] The image and video analysis model uses the generated features to scan the entire video and search for relevant scenes.
[1101] 4. Emotion Engine:
[1102] An algorithm for recognizing the user's emotions and searching and selecting scenes according to their emotional state.
[1103] Facial expression and voice data are analyzed to evaluate the user's emotional state.
[1104] Example of operation
[1105] For example, if a user searches for a scene of a specific player playing in a sports broadcast in 2023, they would proceed as follows:
[1106] 1. Upload video and enter text:
[1107] A user uses his / her own terminal to upload a video file of a live sports broadcast to a server.
[1108] Enter a keyword, such as "a specific player's play," into the text entry field.
[1109] While the user is watching the video, facial expression and voice data is collected by the device's sensors.
[1110] 2. Receiving input data and initial processing:
[1111] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data.
[1112] The video file is temporarily saved, and text and emotion data are passed to the generative AI model and emotion engine.
[1113] 3. Launch and analyze the generative AI model:
[1114] The server launches the generative AI model, and the natural language processing model analyzes the text data. It performs keyword analysis based on the keyword "play by a specific player."
[1115] The image and video analysis model scans the entire video based on the generated features and searches for relevant scenes.
[1116] The emotion engine analyzes facial expression and voice data to determine the user's emotional state.
[1117] 4. Scene detection and extraction:
[1118] The server detects scenes based on the output data from the generative AI model and the emotion analysis results from the emotion engine.
[1119] If the emotion engine rates the user's happiness highly, it sets the priority of that scene high.
[1120] 5. Highlight video generation and delivery:
[1121] The server extracts the detected scene using a video editing library (e.g., FFmpeg or OpenCV) based on the time information of the detected scene.
[1122] Multiple extracted scenes are combined and edited into a highlight video.
[1123] The completed highlight video is saved, and a URL for the saved location is generated and provided to the user.
[1124] In this way, by combining emotion engines, the present invention makes it possible to select scenes according to the individual emotional state of the user, and to generate attractive content with high accuracy.
[1125] Prompt Sentence Examples
[1126] It searches for and extracts specific scenes based on text data entered by the user, such as "a specific player's play from a sports broadcast in 2023."
[1127] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1128] Step 1: Upload video and enter text
[1129] Users select video files from a dedicated application or browser on their own devices and upload them to the server. Users enter keywords or descriptions in a text input field to specify specific scenes. Emotional data such as facial expressions and voice data are also collected while the user is watching the video.
[1130] Input: Video file, text data (e.g., "A specific player's play"), emotional data (facial expressions, voice)
[1131] Output: The video file, text data, and emotion data sent by the user are uploaded to the server.
[1132] Specific operation: The user selects a video file and clicks the upload button. Next, the user enters a keyword specifying a specific scene in the text input field and clicks the send button. While watching the video, the user's facial expression data and voice data are collected by the device's sensors.
[1133] Step 2: Receiving input data and initial processing
[1134] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, and the text data and emotion data are prepared for analysis by the generative AI model and emotion engine.
[1135] Input: Video files uploaded to the server, text data, emotion data
[1136] Output: Temporarily saved video files, analyzable text data and emotion data
[1137] Specific operation: The server analyzes the received HTTP request and saves the video file in the specified directory. Next, it stores the text data and emotion data in memory and prepares for analysis.
[1138] Step 3: Launch and analyze the generative AI model
[1139] The server launches the generative AI model. Text data is input into the natural language processing model, which performs keyword and semantic analysis. Based on the analyzed features, the image and video analysis model scans the entire video to search for relevant scenes. At the same time, the emotion engine analyzes the user's emotional data and determines their emotional state.
[1140] Input: text data, video files, emotion data
[1141] Output: Keyword analysis results, candidate scenes, and user's emotional state
[1142] How it works: The server analyzes the text data entered into the natural language processing model and analyzes the keyword "a specific player's play." Next, based on the generated features, the image and video analysis model scans the entire video to search for specific scenes. At the same time, the emotion engine analyzes facial expression and voice data to evaluate the user's emotions.
[1143] Step 4: Scene detection and extraction
[1144] The server detects scenes based on the output data of the generative AI model and the emotion analysis results of the emotion engine. The emotion engine selects scenes that evoke a specific emotion (e.g., happiness) and sets a priority.
[1145] Input: Analysis results of the generative AI model, analysis results of the emotion engine
[1146] Output: Time information of detected scenes, prioritized scenes
[1147] Specific operation: The server identifies the start and end times of scenes corresponding to Yanagida's at-bats, sets priorities based on the results of emotion analysis, and adds high-priority scenes to a list.
[1148] Step 5: Generate and serve highlight videos
[1149] The server uses a video editing library (e.g., FFmpeg or OpenCV) to extract the relevant parts based on the time information of the detected scenes. If necessary, multiple scenes are combined and edited into a highlight video. The completed highlight video is saved, and a URL for the saved location is generated and provided to the user.
[1150] Input: Time information of detected scenes, prioritized scenes
[1151] Output: Saved highlight video, access URL to the highlight video
[1152] Specific operation: The server uses FFmpeg commands to extract specific scenes from a video file, combines multiple scenes into a single highlight video, and saves the generated highlight video to storage. Finally, it generates a URL for the highlight video and notifies the user.
[1153] (Application example 2)
[1154] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1155] In conventional highlight video generation systems, the detection and extraction of specific scenes is often performed solely through analysis of text data, without taking into account the user's emotions or preferences, resulting in the generated highlight videos not necessarily meeting the user's expectations. Furthermore, personalized content that reflects the user's emotions is in demand, particularly in entertainment applications, but few systems have this functionality.
[1156] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1157] In this invention, the server includes means for uploading video files, means for inputting text data specifying specific scenes, means for analyzing the video files based on the input text data using a generative AI model to detect specific scenes, means for extracting the detected scenes and generating a highlight video, means for collecting user emotion data, means for determining the importance of scenes based on the emotion data, and means for providing the generated highlight video to the user, thereby enabling the generation of a personalized highlight video that takes user emotions into consideration.
[1158] A "video file" is a digital file containing visual and audio information.
[1159] "Uploading" is the act of transferring data from a user's terminal to a remote computer such as a server.
[1160] "Text data" is digital data that contains written information.
[1161] A "generative AI model" is a set of artificial intelligence algorithms designed to perform natural language processing and image and video analysis.
[1162] An "emotion engine" is an algorithm that analyzes the user's emotional state and performs processing based on that data.
[1163] A "highlight video" is a shortened version of a video that has been edited to extract specific scenes.
[1164] "User" means a person or organization that uses the system.
[1165] A "natural language processing model" is an artificial intelligence model for understanding and analyzing text data.
[1166] An "image and video analysis model" is an artificial intelligence model for analyzing and recognizing the content of images and videos.
[1167] "Facial recognition" is a technology that detects and identifies human faces from images and videos.
[1168] "Behavior recognition" is a technology that detects and identifies human actions and movements from images and videos.
[1169] "Emotion data" is digital data that expresses the user's emotional state using numerical values and categories.
[1170] "Decision making" is the process of reaching a conclusion based on data.
[1171] System Overview
[1172] This invention relates to a system that recognizes a user's emotions, automatically searches and extracts specific scenes from videos based on those emotions, and generates highlight videos. The system consists of the following main components:
[1173] 1. Terminal: A device where a user uploads video files and enters text data specifying specific scenes. Typically, this is a smartphone or tablet.
[1174] 2. Server: A computer system that uses a generative AI model to analyze video files, detect and extract specific scenes, and generate a highlight video.
[1175] 3. Generative AI model: A set of algorithms, including natural language processing models and image and video analysis models, used to search and extract specific scenes.
[1176] 4. Emotion engine: An algorithm that analyzes the user's emotional data and determines the importance of a scene based on that data.
[1177] System Operation
[1178] Upload video and enter keywords
[1179] Users upload video files to the server via a smartphone app. They input keywords and descriptions in text format to specify specific scenes. While the user is watching the video, the camera and microphone are used to collect facial expressions and voice data.
[1180] Data reception and initial processing
[1181] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily saved, and the text data is input into a natural language processing model. The emotion data is sent to the emotion engine.
[1182] Analysis of generative AI models
[1183] The server launches a generative AI model to analyze the text data. A natural language processing model (e.g., TensorFlow model) performs keyword and semantic analysis and generates features. Next, an image and video analysis model scans the entire video based on the generated features to search for relevant scenes. At the same time, an emotion engine analyzes the user's emotional data and determines their emotional state (e.g., Microsoft Azure Emotion API).
[1184] Scene detection and extraction
[1185] Based on the analysis results of the generative AI model and emotion engine, the server detects specific scenes. It determines the importance of each scene based on the emotion data and extracts high-priority scenes. It then uses a video editing library (e.g., FFmpeg) to extract the detected scenes and generate a highlight video.
[1186] Highlight video provided
[1187] The generated highlight video is stored on the server and a generated link is provided to the user, who can use the link to view and share the highlight video.
[1188] Specific examples
[1189] For example, if a user wants to search for "scenes of a specific player from a sports broadcast in 2023," the process would proceed as follows: The user uploads a video file of the sports broadcast to the server using their device and enters the keyword "scenes of a specific player." The emotion engine also collects facial expressions and voice data while the user is watching the video. The server inputs the text data and video file into the generative AI model, and performs keyword analysis using a natural language processing model. The image and video analysis model then scans the entire video to search for and detect relevant scenes. At the same time, the emotion engine analyzes the user's emotional data and selects scenes with high importance. The identified scenes are then extracted and edited to create a highlight video. Finally, the user can access the highlight video using the generated link.
[1190] Examples of prompt statements
[1191] "Look for scenes that include a specific player."
[1192] "Create important highlights based on the emotional data of this scene."
[1193] This system enables the generation of personalized highlight videos that reflect the user's emotions.
[1194] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1195] Step 1:
[1196] A user uploads a video file using a smartphone app and inputs text data (keywords and descriptions) to specify a specific scene. While the user is watching the video, the device uses a camera and microphone to collect facial expressions and voice data. The collected data is then sent to a server.
[1197] Input: Video files, text data, facial expressions and audio data
[1198] Output: Send data to the server
[1199] Specific behavior:
[1200] The application accepts user operations and displays a UI that allows the user to select a video file.
[1201] Provide a text field for entering keywords and descriptions.
[1202] It asks for permission to use the camera and microphone, and if the user agrees, it begins collecting data.
[1203] Step 2:
[1204] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, the text data is sent to a natural language processing model (e.g., TensorFlow model), and the emotion data is sent to an emotion engine (e.g., Microsoft Azure Emotion API).
[1205] Input: HTTP request
[1206] Output: Video file, text data, emotion data extraction
[1207] Specific behavior:
[1208] The server receives the HTTP request and parses the video file and text data.
[1209] Temporarily save video files to a file system or cloud storage.
[1210] The text data and emotion data are sent to the corresponding analysis units.
[1211] Step 3:
[1212] The generative AI model analyzes the text data. The natural language processing model performs keyword and semantic analysis and generates feature vectors. Based on these feature vectors, the image and video analysis model scans the entire video and searches for relevant scenes.
[1213] Input: Text data
[1214] Output: Features, list of detected scenes
[1215] Specific behavior:
[1216] A natural language processing model receives the text data and analyzes it for keywords and meaning.
[1217] Features are generated and the data is passed to an image / video analysis model.
[1218] The model scans the video file and lists the start and end times of the relevant scenes.
[1219] Step 4:
[1220] The emotion engine analyzes the user's emotion data and determines the emotional state the user is in. Based on the analysis results, the importance of the scene is determined.
[1221] Input: Emotion data
[1222] Output: Emotional state analysis result, scene importance judgment
[1223] Specific behavior:
[1224] The emotion engine analyzes the user's facial expressions and voice data to detect emotional states such as happiness, surprise, and excitement.
[1225] Emotional data is collected for each scene and the importance of that scene is calculated.
[1226] Step 5:
[1227] The server detects and extracts specific scenes based on the analysis results of the generated AI model and emotion engine, and uses a video editing library (e.g., FFmpeg) to extract the detected scenes and generate a highlight video.
[1228] Input: List of scenes, emotional state analysis results
[1229] Output: Highlight video
[1230] Specific behavior:
[1231] The list of detected scenes is compared with the emotion data to select the scenes to extract.
[1232] Using a video editing library (FFmpeg), selected scenes are extracted from the video file and combined and edited into a highlight video.
[1233] Step 6:
[1234] The server stores the generated highlight video and generates a URL for the storage location. The URL is provided to the user, allowing the user to access the generated highlight video.
[1235] Input: highlight video
[1236] Output: Highlight video URL
[1237] Specific behavior:
[1238] Save highlight videos to a file system or cloud storage.
[1239] Generate a URL for the storage location and provide it via a notification mechanism to the user (e.g., email or in-app notification).
[1240] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1241] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1242] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1243] [Fourth embodiment]
[1244] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1245] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1246] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1247] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1248] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1249] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1250] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1251] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1252] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1253] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1254] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1255] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1256] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1257] The present invention is applied to a system that automatically searches for specific scenes from a video file and generates a highlight video. The system mainly includes a user, a terminal, and a server.
[1258] Basic system configuration
[1259] The system according to the present invention comprises the following main components:
[1260] 1. User's device: A device where video files are uploaded and text data specifying specific scenes is entered.
[1261] 2. Server: A computer system that uses a generative AI model to analyze video and extract scenes, then generates a highlight video.
[1262] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models to search and extract specific scenes.
[1263] Program processing flow
[1264] Upload video and enter text
[1265] Users upload video files to the server through a browser or a dedicated application. They also enter keywords and descriptions to specify specific scenes in a text input field. The device sends this data to the server as a single HTTP request.
[1266] Receiving and initial processing of input data
[1267] The server analyzes the received HTTP request and extracts the video file and text data. The video file is temporarily stored, and the text data is used for the next stage of analysis.
[1268] Launching and analyzing generative AI models
[1269] The server launches the generative AI model. First, text data is input into the natural language processing model, and keyword and semantic analysis is performed. As a specific example, the keyword "Yanagida's turn at bat" is analyzed and its features are generated.
[1270] Next, based on the generated features, the image and video analysis model scans the entire video to search for relevant scenes, and uses facial and behavioral recognition algorithms to identify the start and end times of scenes that match the analyzed features.
[1271] Scene detection and extraction
[1272] The server detects scenes based on the output data from the generative AI model. Using the time information of the identified scenes, it extracts the scenes using a video editing library (e.g., FFmpeg or OpenCV). The extracted scenes are temporarily saved as individual clips.
[1273] Highlight video generation and provision
[1274] The server combines the extracted scenes and edits them into a highlight video. The completed highlight video is saved and a URL for the saved location is generated. This URL is provided to the user, who can then view or download the generated highlight video via the link.
[1275] Specific examples
[1276] For example, if a user is searching for a scene of "Yanagida's turn at bat in a baseball broadcast in 2023," the process would be as follows: First, the user uses their device to upload a video file of the baseball broadcast to the server and enters the keyword "Yanagida's turn at bat." The server inputs the text data and video file into the generative AI model, and performs keyword analysis using a natural language processing model.
[1277] Next, an image and video analysis model scans the entire video to search for and detect scenes that correspond to Yanagi's at-bat. The identified scenes are extracted and edited to create a highlight video. Finally, users can access the generated highlight video using the provided link.
[1278] In this way, the system according to the present invention enables the user to easily search for a particular scene and realizes the process of efficiently generating a highlight video.
[1279] The processing flow will be explained below.
[1280] Step 1:
[1281] Users select video files from their own devices and upload them to the server via a browser or dedicated application.
[1282] Step 2:
[1283] The user inputs text data specifying a particular scene into the input field, for example, specifying the keyword "Yanagida's turn at bat."
[1284] Step 3:
[1285] The terminal sends the video file and text data to the server as a single HTTP request.
[1286] Step 4:
[1287] The server analyzes the received HTTP request and extracts the video file and text data. The video file is temporarily stored, and the text data is used to analyze the generative AI model.
[1288] Step 5:
[1289] The server launches the generative AI model. First, it inputs text data into the natural language processing model, analyzes the keyword "Yanagida's turn at bat," and generates its features.
[1290] Step 6:
[1291] The server inputs the generated features into an image / video analysis model, which then scans the entire video to find scenes that match the features.
[1292] Step 7:
[1293] The image and video analysis model detects the start and end times of scenes corresponding to "Yanagita's turn at bat" from the video and returns them to the server.
[1294] Step 8:
[1295] The server extracts the scene from the video file using a video editing library (e.g., FFmpeg or OpenCV) based on the time information of the detected scene.
[1296] Step 9:
[1297] The server temporarily stores the extracted scenes as one clip.
[1298] Step 10:
[1299] The server combines a plurality of extracted scenes as necessary and edits them into a highlight video.
[1300] Step 11:
[1301] The server saves the completed highlight video and generates a URL for the location where it is saved.
[1302] Step 12:
[1303] The server sends the generated URL to the terminal as an HTTP response.
[1304] Step 13:
[1305] The terminal displays the received URL to the user.
[1306] Step 14:
[1307] The user can click on the displayed link to view or download the generated highlight video.
[1308] Example 1
[1309] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1310] In modern video media, the task of efficiently extracting specific scenes from long video data and generating the highlight footage desired by users is extremely time-consuming and labor-intensive. Manually editing large amounts of video data is particularly impractical, creating a demand for automated systems. However, existing systems lack the technology for highly accurate scene detection and highlight generation. To address this issue, a system that achieves more accurate scene detection and highlight generation is needed.
[1311] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1312] In this invention, the server includes means for uploading video data, means for inputting text data specifying specific scenes, means for analyzing the video data based on the input text data using a generative AI model to detect the specific scenes, means for extracting the detected scenes and generating a highlight video, and means for providing the generated highlight video to the user. This makes it possible to automatically extract scenes specified by the user with high accuracy and easily generate a highlight video.
[1313] "Video data" refers to video files stored in digital format and has the attributes of a media file.
[1314] "Text data" is character string information that the user inputs to specify a particular scene.
[1315] A "generative AI model" is a model that uses artificial intelligence technology and includes a series of algorithms to analyze video data based on text data and detect specific scenes.
[1316] A "natural language processing model" is part of a generative AI model and is an algorithm for analyzing the meaning of input text data and extracting keywords.
[1317] The "image and video analysis model" is part of the generative AI model and is an algorithm for searching and detecting specific scenes within video data.
[1318] A "person recognition algorithm" is a technology for identifying specific people within video data.
[1319] "Movement recognition algorithms" are techniques for identifying specific movements or actions within video data.
[1320] "Cutting out" is an operation for extracting a specific portion of video data based on the start and end times of a detected scene.
[1321] A "highlight video" is a short video file created by combining specific scenes, and is a format that allows users to efficiently view scenes they want to focus on.
[1322] A "user" is an individual or group that operates the system to upload video data and specify specific scenes.
[1323] The present invention is applied to a system that automatically searches for specific scenes from video data and generates highlight videos. The system mainly includes a user, a terminal, and a server.
[1324] Basic system configuration
[1325] The system according to the present invention comprises the following main components:
[1326] 1. Terminal: A device where a user uploads video data and inputs text data specifying a particular scene.
[1327] 2. Server: A computer system that uses a generative AI model to analyze video data, extract scenes, and generate highlight footage.
[1328] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models that search for and extract specific scenes.
[1329] Example of operation
[1330] For example, if a user requests "a specific player's play from a sports broadcast in 2023," the process would proceed as follows: The user uses their device to upload a video file of the sports broadcast to the server and enters the keyword "a specific player's play." The server inputs the text data and video data into the generative AI model, and performs keyword analysis using a natural language processing model.
[1331] Next, based on the generated features, the image and video analysis model scans the entire video data to search for relevant scenes, and uses facial and behavior recognition algorithms to identify the start and end times of scenes that match the analyzed features.
[1332] Once the scenes are identified, the server uses a video editing library (e.g., FFmpeg or OpenCV) to extract the segments, which are then temporarily saved as individual clips.
[1333] Finally, the server combines the extracted scenes and edits them into a highlight video. The completed highlight video is saved on the server, and a URL for the saved location is generated. This URL is provided to users, who can then view or download the created highlight video via the link.
[1334] Hardware and software used
[1335] Hardware: User's device (PC, smartphone, etc.), server
[1336] Software: Generative AI models, video editing libraries (FFmpeg and OpenCV), natural language processing models (BERT, GPT, etc.), image and video analysis models (YOLO, etc.)
[1337] Prompt Sentence Examples
[1338] "Automatically generate highlight footage including a specific player's play from a sports broadcast in 2023."
[1339] By using this system, users can efficiently extract specific scenes from long video data and generate highlight videos in a short amount of time.
[1340] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1341] Step 1:
[1342] The user uploads video data using a device and enters text data to specify a specific scene. Specifically, the user selects video data through a browser or a dedicated application and enters keywords such as "a specific player's play." The device then sends this data to the server as a single HTTP request.
[1343] Input: Video data, text data
[1344] Output: HTTP request
[1345] Step 2:
[1346] The server analyzes the received HTTP request and extracts video and text data. The server temporarily stores the video data, and the text data is used in the next analysis stage. A parsing process is performed to analyze and extract data from the HTTP request. Keywords such as "plays by a specific player" are extracted and stored as input data.
[1347] Input: HTTP request
[1348] Output: Video data, text data
[1349] Step 3:
[1350] The server launches the generative AI model. First, the text data is input into the natural language processing model, which performs keyword analysis and semantic analysis. As a specific example, the natural language processing model analyzes the "play of a specific player" and generates its features. These features become indicators for searching the video data.
[1351] Input: Text data
[1352] Output: Text features
[1353] Step 4:
[1354] The server launches the image and video analysis model, scans the video data based on the generated text features, and searches for relevant scenes. It then uses facial and behavior recognition algorithms to identify the start and end times of specific scenes. Specifically, it analyzes the video data frame by frame to find scenes that match the features.
[1355] Input: Video data, text features
[1356] Output: Scene start and end times
[1357] Step 5:
[1358] The server uses a video editing library (e.g., FFmpeg, OpenCV) to extract the identified scene based on its time information. The extracted scene is temporarily saved as an individual clip. Specifically, the server uses a library such as FFmpeg to extract video within a specified time range.
[1359] Input: Video data, scene start and end times
[1360] Output: Video clip
[1361] Step 6:
[1362] The server combines multiple scenes and edits them into a highlight video. The completed highlight video is saved on the server and a URL for its location is generated. Specifically, an editing library such as FFmpeg is used to combine the individual clips into a single video file.
[1363] Input: Video clip
[1364] Output: highlight video file, viewing URL
[1365] Step 7:
[1366] The user can view or download the generated highlight video using the provided viewing URL. Specifically, the user opens the URL in a browser and plays or saves the video.
[1367] Input: Viewing URL
[1368] Output: Played video or downloaded video file
[1369] (Application example 1)
[1370] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1371] Conventional video search systems have had the problem that it takes a lot of time and effort to find a specific scene. Manually editing a highlight video is also cumbersome, placing a significant burden on the user. Furthermore, there has been a lack of means for efficiently viewing the generated highlight video, making it difficult to improve the user experience. The present invention aims to solve these problems by providing a system that enables users to easily search for specific scenes, generate a highlight video, and efficiently view it.
[1372] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1373] In this invention, the server includes means for uploading video files, means for inputting text data specifying specific scenes, means for analyzing the video files based on the input text data using a generative AI model and detecting specific scenes, means for extracting the detected scenes and generating a highlight video, means for using cloud storage to provide the generated highlight video to a user, and means for providing the generated highlight video by streaming, thereby enabling a user to efficiently search for specific scenes with little effort, generate a highlight video, and further watch the highlight video by streaming.
[1374] "Means for uploading video files" refers to a function that allows a user to send video data from their own device to a server and save it.
[1375] "Means for inputting text data specifying a particular scene" refers to an input interface that a user uses to specify a scene to be searched for using keywords or a description.
[1376] "Means of using a generative AI model to analyze video files based on input text data and detect specific scenes" refers to the function in which a generative AI model analyzes text data and identifies the relevant scenes in the video based on the results.
[1377] The "means for extracting detected scenes and generating a highlight video" refers to a function for cutting out identified scenes from the original video and editing them to create a highlight video.
[1378] The "means of using cloud storage to provide the generated highlight video to the user" refers to a function for storing the generated highlight video on the cloud and enabling the user to access it.
[1379] The term "means for providing the generated highlight video by streaming" refers to a function of transmitting the video via the Internet so that the user can view the highlight video generated in real time.
[1380] Basic system configuration
[1381] The system according to the present invention comprises the following main components:
[1382] 1. User's device: A device on which a video file is uploaded using a smartphone app and text data specifying a specific scene is entered.
[1383] 2. Server: A computer system that performs video analysis, scene extraction, and provides the generated highlight video to users. This server runs on Amazon AWS or other cloud platforms.
[1384] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models that search for and extract specific scenes.
[1385] Upload video and enter text
[1386] A user uploads a video file to a server through a smartphone app. Then, they enter keywords and descriptions to specify specific scenes in a text input field. For example, they enter the keyword "live performance of a popular artist." This input data is sent from the application to the server as a single HTTP request.
[1387] Receiving and initial processing of input data
[1388] The server analyzes the received HTTP request and extracts the video file and text data. The video file is temporarily stored in cloud storage such as Amazon S3, and the text data is used for the next stage of analysis.
[1389] Launching and analyzing generative AI models
[1390] The server launches a generative AI model. First, it inputs text data into a natural language processing (NLP) model (e.g., GPT-4) to perform keyword and semantic analysis. This operation identifies the content that the input keywords contain. Next, it uses an image / video analysis model (e.g., OpenCV or YOLO) to scan the entire video based on the analyzed keywords and search for relevant scenes.
[1391] Scene detection and extraction
[1392] Using the time information of the identified scene, the part is extracted using a video editing library (e.g., FFmpeg). The detected scene is temporarily saved as a clip.
[1393] Highlight video generation and provision
[1394] The server combines multiple scenes and edits them into a highlight video. The generated highlight video is saved to a server such as Amazon S3, and a URL for the saved location is generated. Users can view or download the generated highlight video from the app via this URL. The highlight video is also provided in real time via a streaming service.
[1395] Specific examples
[1396] For example, if a user searches for "live performances by popular artists," the process would proceed as follows: The user uses their device to upload a video file of the live performance to the server and enters the keyword "live performances by popular artists." The server uses a generative AI model to analyze the text data and video file, and a natural language processing model to perform keyword analysis. Next, an image and video analysis model scans the entire video to search for relevant scenes, and identified scenes are extracted as individual clips. These clips are then edited to create a highlight video, which is then provided to the user.
[1397] Prompt Sentence Examples
[1398] "Search for live performance scenes of popular artists and generate highlight videos."
[1399] In this way, the system according to the present invention realizes a process that allows a user to easily search for a particular scene and efficiently generate and view a highlight video.
[1400] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1401] Step 1:
[1402] A user uploads a video file to a server through a smartphone app. The input is a video file from the user's device, and the output is a video file stored on the server. Specifically, the user uses the app's "video upload" function to send a video file via an HTTP request. The server receives the file and stores it in cloud storage such as Amazon S3.
[1403] Step 2:
[1404] Within the same application, the user enters text data to specify a particular scene. The input is keywords and descriptions related to the particular scene, and the output is text data sent to the server. Specifically, the user enters keywords in the text field and presses the "Search" button. The text data is sent to the server along with the video file via an HTTP request.
[1405] Step 3:
[1406] The server analyzes the received video file and text data. The input is the video file and text data, and the output is the text analysis results. Specifically, the server retrieves the video file from Amazon S3 and simultaneously supplies the text data to a natural language processing (NLP) model (e.g., GPT-4). The NLP model performs keyword analysis and semantic analysis to generate the necessary features.
[1407] Step 4:
[1408] The image and video analysis part of the generative AI model is activated, and it scans the entire video based on the analyzed features to search for specific scenes. The input is the analysis results from the NLP model and the video file, and the output is the time information of the detected scene. Specifically, the generative AI model (e.g., OpenCV or YOLO) analyzes the entire video frame by frame and identifies the start and end times of the relevant scenes.
[1409] Step 5:
[1410] The server extracts the detected scenes and generates highlight clips. The input is the scene's time information and the original video file, and the output is the extracted highlight clip. Specifically, the server uses a video editing library (e.g., FFmpeg) to extract a specific scene from its start time to its end time and temporarily save it as an individual clip.
[1411] Step 6:
[1412] The server then combines the extracted scenes into a single highlight video. The input is multiple highlight clips, and the output is the completed highlight video. Specifically, the server uses FFmpeg again to combine the temporarily saved clips into a single video file.
[1413] Step 7:
[1414] The server saves the generated highlight video in cloud storage and generates a URL for the storage location. The input is the completed highlight video, and the output is the video URL. Specifically, the server uploads the completed highlight video to Amazon S3 and generates an access URL.
[1415] Step 8:
[1416] The user uses the provided URL to watch or download the generated highlight video. The input is the video URL, and the output is the highlight video played on the user's device. Specifically, the user clicks the URL link in the smartphone app and streams or downloads the video for viewing.
[1417] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1418] This invention is applied to a system that automatically searches for specific scenes from video files and generates highlight videos by combining an emotion engine that recognizes the user's emotions. The system selects scenes that correspond to the user's emotional state.
[1419] Basic system configuration
[1420] The system according to the present invention comprises the following main components:
[1421] 1. User's device: A device where video files are uploaded and text data specifying specific scenes is entered.
[1422] 2. Server: A computer system that uses a generative AI model to analyze videos and extract scenes, and combines it with the user's emotion engine to generate a highlight video.
[1423] 3. Generative AI model: A set of algorithms including natural language processing models and image and video analysis models to search and extract specific scenes.
[1424] 4. Emotion Engine: An algorithm to recognize the user's emotions and search and select scenes according to their emotional state.
[1425] Program processing flow
[1426] Upload video and enter text
[1427] Users upload video files to the server through a browser or a dedicated application. They also enter keywords or descriptions to identify specific scenes in a text input field. The system also collects data to recognize the user's emotions. For example, it records facial expression data and voice input while the user is watching a specific video scene.
[1428] Receiving and initial processing of input data
[1429] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, and the text data and emotion data are used for analysis by the generative AI model and emotion engine.
[1430] Launching and analyzing generative AI models
[1431] The server launches the generative AI model. First, text data is input into the natural language processing model, and keyword and semantic analysis is performed. As a specific example, the keyword "Yanagida's turn at bat" is analyzed and its features are generated.
[1432] Next, based on the generated features, the image and video analysis model scans the entire video to find relevant scenes. In addition, the emotion engine analyzes the user's emotional data and determines their emotional state.
[1433] Scene detection and extraction
[1434] The server detects scenes based on the output data from the generative AI model and the emotion analysis results from the emotion engine. For example, if the emotion engine recognizes a high level of happiness in the user, it determines that the scene is particularly important and sets a high priority for it.
[1435] Highlight video generation and provision
[1436] The server uses the time information of the detected scenes to extract those parts using a video editing library (e.g., FFmpeg or OpenCV). The extracted scenes are temporarily saved as a single clip. Multiple extracted scenes are then combined as needed to create a highlight video. The completed highlight video is saved, and a URL for its save location is generated.
[1437] Reflecting user reactions based on emotional data
[1438] The system accumulates user emotion data through the emotion engine, which can be used to improve the accuracy of future analysis results. Also, by reflecting the user's emotions toward specific scenes, the system can provide more accurate highlight videos.
[1439] Specific examples
[1440] For example, if a user is searching for a scene of "Yanagida's turn at bat in a baseball broadcast in 2023," the process would be as follows. First, the user uses their device to upload a video file of the baseball broadcast to the server and enters the keyword "Yanagida's turn at bat." The emotion engine then collects facial expressions and voice data while the user is watching the video. The server then inputs the text data and video file into a generative AI model, and performs keyword analysis using a natural language processing model.
[1441] Next, an image and video analysis model scans the entire video to search for and detect scenes that correspond to Yanagi's turn at bat. At the same time, an emotion engine analyzes the user's emotional data and prioritizes scenes in which the user expresses happiness. The identified scenes are extracted and edited to create a highlight video. Finally, the user can access the generated highlight video using the provided link.
[1442] In this way, by combining emotion engines, the present invention enables scene selection according to the individual emotional state of the user, and generates attractive content with high accuracy.
[1443] The processing flow will be explained below.
[1444] Step 1:
[1445] Users select video files from their own devices and upload them to the server via a browser or dedicated application.
[1446] Step 2:
[1447] The user inputs text data specifying a particular scene into the input field, for example, specifying the keyword "Yanagida's turn at bat."
[1448] Step 3:
[1449] The user gives permission to the system to record emotional data (facial expressions and voice) while watching the video.
[1450] Step 4:
[1451] The terminal transmits the video file, text data, and emotion data to the server as a single HTTP request.
[1452] Step 5:
[1453] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, and the text data and emotion data are used for analysis by the generative AI model and emotion engine.
[1454] Step 6:
[1455] The server launches the generative AI model. Text data is input into the natural language processing model, and semantic analysis of the text is performed. For example, the keyword "Yanagida's turn at bat" is analyzed and features are generated.
[1456] Step 7:
[1457] The server runs an image / video analysis model based on the generated features and searches the entire video for corresponding scenes. The image / video analysis model detects specific scenes using facial and behavior recognition algorithms.
[1458] Step 8:
[1459] The emotion engine analyzes the user's emotion data (facial expression data and voice data) to determine the user's emotional state, for example, identifying emotion categories such as happiness, excitement, and anxiety.
[1460] Step 9:
[1461] Based on the analysis results of the emotion engine, the server evaluates scenes in the video based on the emotion data and feature values, and sets priorities. Scenes in which the user expressed a favorable emotion are given a higher priority.
[1462] Step 10:
[1463] Based on the evaluation results, the server extracts specific scenes from the video file by identifying the start and end times of the scenes using a video editing library (such as FFmpeg or OpenCV) and extracting those parts.
[1464] Step 11:
[1465] The server temporarily stores the extracted scenes as one clip.
[1466] Step 12:
[1467] The server combines multiple extracted scenes as needed and edits them into a highlight video based on the user's emotional state.
[1468] Step 13:
[1469] The server saves the completed highlight video and generates a URL for the location where it is saved.
[1470] Step 14:
[1471] The server sends the generated URL to the terminal as an HTTP response.
[1472] Step 15:
[1473] The terminal displays the received URL to the user.
[1474] Step 16:
[1475] The user can click on the displayed link to view or download the generated highlight video.
[1476] Example 2
[1477] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1478] Conventional systems were unable to consider the user's subjective emotions when searching for specific scenes within a video file and generating a highlight video. This made it difficult to appropriately select the scenes that appealed to each individual user. Furthermore, there was a lack of technology for performing highly accurate analysis using user emotion data. This has led to a demand for improved user satisfaction.
[1479] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1480] In this invention, the server includes means for uploading video files, means for inputting text data specifying specific scenes, means for analyzing the video files based on the input text data using a generative AI model and detecting specific scenes, means for collecting user emotion data, means for analyzing the emotion data and selecting scenes based on the user's emotional state, means for extracting the detected scenes and generating a highlight video, and means for providing the generated highlight video to the user. This enables scene selection that takes user emotion into consideration and highly accurate highlight video generation.
[1481] 1. "Means for uploading video files" means the technical means that allows a user to transmit video files to a server via the Internet.
[1482] 2. "Means for inputting text data specifying a specific scene" means a technical means for providing an interface for a user to input text data such as keywords or descriptions in order to specify a specific video scene.
[1483] 3. "Generative AI model" refers to a set of algorithms that analyze input data and process it for a specific purpose, including, specifically, natural language processing models and image and video analysis models.
[1484] 4. "Natural language processing model" refers to an algorithm for analyzing text data and performing keyword analysis and semantic analysis.
[1485] 5. "Image / Video Analysis Model" refers to an algorithm that analyzes scenes in a video file and detects specific actions or objects.
[1486] 6. "Means for collecting user emotional data" refers to technical means for detecting a user's facial expressions, voice, actions, etc., and collecting emotional data based on them.
[1487] 7. "Means for analyzing emotional data and selecting scenes based on the user's emotional state" means technical means for determining the user's emotional state using collected emotional data and selecting important scenes based on that emotional state.
[1488] 8. "Means for extracting detected scenes and generating a highlight video" refers to the technical means of using a video editing library to extract that portion based on the time information of the detected scenes and edit it into a single highlight video.
[1489] 9. "Means for providing the generated highlight video to the user" refers to the technical means for saving the completed highlight video, generating a URL for the location where the video is saved, and providing the URL to the user.
[1490] In this way, by combining each of these means, it is possible to generate a highlight video that reflects the user's emotions.
[1491] The present invention relates to a system that automatically searches for specific scenes from video files and generates highlight videos by combining an emotion engine that recognizes the user's emotions. This system selects scenes that correspond to the user's emotional state.
[1492] Basic system configuration
[1493] The system according to the present invention comprises the following main components:
[1494] 1. On the user's device:
[1495] A user uses a device (such as a computer or smartphone) to upload a video file.
[1496] The device sends the video file to the server via a browser or a dedicated application.
[1497] A text entry field is used to enter keywords or descriptions to specify a particular scene.
[1498] It is also equipped with sensors to collect user emotional data, such as facial expression data and voice data.
[1499] 2. Server:
[1500] The server uses a generative AI model to analyze the video and extract scenes.
[1501] It analyzes received HTTP requests and extracts and saves video files, text data, and emotional data.
[1502] The hardware used is a computer system equipped with a high-performance processor and sufficient memory.
[1503] 3. Generative AI Model:
[1504] It consists of a set of algorithms including natural language processing models and image and video analysis models.
[1505] The natural language processing model analyzes the input text data and performs keyword analysis and semantic analysis.
[1506] The image and video analysis model uses the generated features to scan the entire video and search for relevant scenes.
[1507] 4. Emotion Engine:
[1508] An algorithm for recognizing the user's emotions and searching and selecting scenes according to their emotional state.
[1509] Facial expression and voice data are analyzed to evaluate the user's emotional state.
[1510] Example of operation
[1511] For example, if a user searches for a scene of a specific player playing in a sports broadcast in 2023, they would proceed as follows:
[1512] 1. Upload video and enter text:
[1513] A user uses his / her own terminal to upload a video file of a live sports broadcast to a server.
[1514] Enter a keyword, such as "a specific player's play," into the text entry field.
[1515] While the user is watching the video, facial expression and voice data is collected by the device's sensors.
[1516] 2. Receiving input data and initial processing:
[1517] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data.
[1518] The video file is temporarily saved, and text and emotion data are passed to the generative AI model and emotion engine.
[1519] 3. Launch and analyze the generative AI model:
[1520] The server launches the generative AI model, and the natural language processing model analyzes the text data. It performs keyword analysis based on the keyword "play by a specific player."
[1521] The image and video analysis model scans the entire video based on the generated features and searches for relevant scenes.
[1522] The emotion engine analyzes facial expression and voice data to determine the user's emotional state.
[1523] 4. Scene detection and extraction:
[1524] The server detects scenes based on the output data from the generative AI model and the emotion analysis results from the emotion engine.
[1525] If the emotion engine rates the user's happiness highly, it sets the priority of that scene high.
[1526] 5. Highlight video generation and delivery:
[1527] The server extracts the detected scene using a video editing library (e.g., FFmpeg or OpenCV) based on the time information of the detected scene.
[1528] Multiple extracted scenes are combined and edited into a highlight video.
[1529] The completed highlight video is saved, and a URL for the saved location is generated and provided to the user.
[1530] In this way, by combining emotion engines, the present invention makes it possible to select scenes according to the individual emotional state of the user, and to generate attractive content with high accuracy.
[1531] Prompt Sentence Examples
[1532] It searches for and extracts specific scenes based on text data entered by the user, such as "a specific player's play from a sports broadcast in 2023."
[1533] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1534] Step 1: Upload video and enter text
[1535] Users select video files from a dedicated application or browser on their own devices and upload them to the server. Users enter keywords or descriptions in a text input field to specify specific scenes. Emotional data such as facial expressions and voice data are also collected while the user is watching the video.
[1536] Input: Video file, text data (e.g., "A specific player's play"), emotional data (facial expressions, voice)
[1537] Output: The video file, text data, and emotion data sent by the user are uploaded to the server.
[1538] Specific operation: The user selects a video file and clicks the upload button. Next, the user enters a keyword specifying a specific scene in the text input field and clicks the send button. While watching the video, the user's facial expression data and voice data are collected by the device's sensors.
[1539] Step 2: Receiving input data and initial processing
[1540] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, and the text data and emotion data are prepared for analysis by the generative AI model and emotion engine.
[1541] Input: Video files uploaded to the server, text data, emotion data
[1542] Output: Temporarily saved video files, analyzable text data and emotion data
[1543] Specific operation: The server analyzes the received HTTP request and saves the video file in the specified directory. Next, it stores the text data and emotion data in memory and prepares for analysis.
[1544] Step 3: Launch and analyze the generative AI model
[1545] The server launches the generative AI model. Text data is input into the natural language processing model, which performs keyword and semantic analysis. Based on the analyzed features, the image and video analysis model scans the entire video to search for relevant scenes. At the same time, the emotion engine analyzes the user's emotional data and determines their emotional state.
[1546] Input: text data, video files, emotion data
[1547] Output: Keyword analysis results, candidate scenes, and user's emotional state
[1548] How it works: The server analyzes the text data entered into the natural language processing model and analyzes the keyword "a specific player's play." Next, based on the generated features, the image and video analysis model scans the entire video to search for specific scenes. At the same time, the emotion engine analyzes facial expression and voice data to evaluate the user's emotions.
[1549] Step 4: Scene detection and extraction
[1550] The server detects scenes based on the output data of the generative AI model and the emotion analysis results of the emotion engine. The emotion engine selects scenes that evoke a specific emotion (e.g., happiness) and sets a priority.
[1551] Input: Analysis results of the generative AI model, analysis results of the emotion engine
[1552] Output: Time information of detected scenes, prioritized scenes
[1553] Specific operation: The server identifies the start and end times of scenes corresponding to Yanagida's at-bats, sets priorities based on the results of emotion analysis, and adds high-priority scenes to a list.
[1554] Step 5: Generate and serve highlight videos
[1555] The server uses a video editing library (e.g., FFmpeg or OpenCV) to extract the relevant parts based on the time information of the detected scenes. If necessary, multiple scenes are combined and edited into a highlight video. The completed highlight video is saved, and a URL for the saved location is generated and provided to the user.
[1556] Input: Time information of detected scenes, prioritized scenes
[1557] Output: Saved highlight video, access URL to the highlight video
[1558] Specific operation: The server uses FFmpeg commands to extract specific scenes from a video file, combines multiple scenes into a single highlight video, and saves the generated highlight video to storage. Finally, it generates a URL for the highlight video and notifies the user.
[1559] (Application example 2)
[1560] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1561] In conventional highlight video generation systems, the detection and extraction of specific scenes is often performed solely through analysis of text data, without taking into account the user's emotions or preferences, resulting in the generated highlight videos not necessarily meeting the user's expectations. Furthermore, personalized content that reflects the user's emotions is in demand, particularly in entertainment applications, but few systems have this functionality.
[1562] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1563] In this invention, the server includes means for uploading video files, means for inputting text data specifying specific scenes, means for analyzing the video files based on the input text data using a generative AI model to detect specific scenes, means for extracting the detected scenes and generating a highlight video, means for collecting user emotion data, means for determining the importance of scenes based on the emotion data, and means for providing the generated highlight video to the user, thereby enabling the generation of a personalized highlight video that takes user emotions into consideration.
[1564] A "video file" is a digital file containing visual and audio information.
[1565] "Uploading" is the act of transferring data from a user's terminal to a remote computer such as a server.
[1566] "Text data" is digital data that contains written information.
[1567] A "generative AI model" is a set of artificial intelligence algorithms designed to perform natural language processing and image and video analysis.
[1568] An "emotion engine" is an algorithm that analyzes the user's emotional state and performs processing based on that data.
[1569] A "highlight video" is a shortened version of a video that has been edited to extract specific scenes.
[1570] "User" means a person or organization that uses the system.
[1571] A "natural language processing model" is an artificial intelligence model for understanding and analyzing text data.
[1572] An "image and video analysis model" is an artificial intelligence model for analyzing and recognizing the content of images and videos.
[1573] "Facial recognition" is a technology that detects and identifies human faces from images and videos.
[1574] "Behavior recognition" is a technology that detects and identifies human actions and movements from images and videos.
[1575] "Emotion data" is digital data that expresses the user's emotional state using numerical values and categories.
[1576] "Decision making" is the process of reaching a conclusion based on data.
[1577] System Overview
[1578] This invention relates to a system that recognizes a user's emotions, automatically searches and extracts specific scenes from videos based on those emotions, and generates highlight videos. The system consists of the following main components:
[1579] 1. Terminal: A device where a user uploads video files and enters text data specifying specific scenes. Typically, this is a smartphone or tablet.
[1580] 2. Server: A computer system that uses a generative AI model to analyze video files, detect and extract specific scenes, and generate a highlight video.
[1581] 3. Generative AI model: A set of algorithms, including natural language processing models and image and video analysis models, used to search and extract specific scenes.
[1582] 4. Emotion engine: An algorithm that analyzes the user's emotional data and determines the importance of a scene based on that data.
[1583] System Operation
[1584] Upload video and enter keywords
[1585] Users upload video files to the server via a smartphone app. They input keywords and descriptions in text format to specify specific scenes. While the user is watching the video, the camera and microphone are used to collect facial expressions and voice data.
[1586] Data reception and initial processing
[1587] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily saved, and the text data is input into a natural language processing model. The emotion data is sent to the emotion engine.
[1588] Analysis of generative AI models
[1589] The server launches a generative AI model to analyze the text data. A natural language processing model (e.g., TensorFlow model) performs keyword and semantic analysis and generates features. Next, an image and video analysis model scans the entire video based on the generated features to search for relevant scenes. At the same time, an emotion engine analyzes the user's emotional data and determines their emotional state (e.g., Microsoft Azure Emotion API).
[1590] Scene detection and extraction
[1591] Based on the analysis results of the generative AI model and emotion engine, the server detects specific scenes. It determines the importance of each scene based on the emotion data and extracts high-priority scenes. It then uses a video editing library (e.g., FFmpeg) to extract the detected scenes and generate a highlight video.
[1592] Highlight video provided
[1593] The generated highlight video is stored on the server and a generated link is provided to the user, who can use the link to view and share the highlight video.
[1594] Specific examples
[1595] For example, if a user wants to search for "scenes of a specific player from a sports broadcast in 2023," the process would proceed as follows: The user uploads a video file of the sports broadcast to the server using their device and enters the keyword "scenes of a specific player." The emotion engine also collects facial expressions and voice data while the user is watching the video. The server inputs the text data and video file into the generative AI model, and performs keyword analysis using a natural language processing model. The image and video analysis model then scans the entire video to search for and detect relevant scenes. At the same time, the emotion engine analyzes the user's emotional data and selects scenes with high importance. The identified scenes are then extracted and edited to create a highlight video. Finally, the user can access the highlight video using the generated link.
[1596] Examples of prompt statements
[1597] "Look for scenes that include a specific player."
[1598] "Create important highlights based on the emotional data of this scene."
[1599] This system enables the generation of personalized highlight videos that reflect the user's emotions.
[1600] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1601] Step 1:
[1602] A user uploads a video file using a smartphone app and inputs text data (keywords and descriptions) to specify a specific scene. While the user is watching the video, the device uses a camera and microphone to collect facial expressions and voice data. The collected data is then sent to a server.
[1603] Input: Video files, text data, facial expressions and audio data
[1604] Output: Send data to the server
[1605] Specific behavior:
[1606] The application accepts user operations and displays a UI that allows the user to select a video file.
[1607] Provide a text field for entering keywords and descriptions.
[1608] It asks for permission to use the camera and microphone, and if the user agrees, it begins collecting data.
[1609] Step 2:
[1610] The server analyzes the received HTTP request and extracts the video file, text data, and emotion data. The video file is temporarily stored, the text data is sent to a natural language processing model (e.g., TensorFlow model), and the emotion data is sent to an emotion engine (e.g., Microsoft Azure Emotion API).
[1611] Input: HTTP request
[1612] Output: Video file, text data, emotion data extraction
[1613] Specific behavior:
[1614] The server receives the HTTP request and parses the video file and text data.
[1615] Temporarily save video files to a file system or cloud storage.
[1616] The text data and emotion data are sent to the corresponding analysis units.
[1617] Step 3:
[1618] The generative AI model analyzes the text data. The natural language processing model performs keyword and semantic analysis and generates feature vectors. Based on these feature vectors, the image and video analysis model scans the entire video and searches for relevant scenes.
[1619] Input: Text data
[1620] Output: Features, list of detected scenes
[1621] Specific behavior:
[1622] A natural language processing model receives the text data and analyzes it for keywords and meaning.
[1623] Features are generated and the data is passed to an image / video analysis model.
[1624] The model scans the video file and lists the start and end times of the relevant scenes.
[1625] Step 4:
[1626] The emotion engine analyzes the user's emotion data and determines the emotional state the user is in. Based on the analysis results, the importance of the scene is determined.
[1627] Input: Emotion data
[1628] Output: Emotional state analysis result, scene importance judgment
[1629] Specific behavior:
[1630] The emotion engine analyzes the user's facial expressions and voice data to detect emotional states such as happiness, surprise, and excitement.
[1631] Emotional data is collected for each scene and the importance of that scene is calculated.
[1632] Step 5:
[1633] The server detects and extracts specific scenes based on the analysis results of the generated AI model and emotion engine, and uses a video editing library (e.g., FFmpeg) to extract the detected scenes and generate a highlight video.
[1634] Input: List of scenes, emotional state analysis results
[1635] Output: Highlight video
[1636] Specific behavior:
[1637] The list of detected scenes is compared with the emotion data to select the scenes to extract.
[1638] Using a video editing library (FFmpeg), selected scenes are extracted from the video file and combined and edited into a highlight video.
[1639] Step 6:
[1640] The server stores the generated highlight video and generates a URL for the storage location. The URL is provided to the user, allowing the user to access the generated highlight video.
[1641] Input: highlight video
[1642] Output: Highlight video URL
[1643] Specific behavior:
[1644] Save highlight videos to a file system or cloud storage.
[1645] Generate a URL for the storage location and provide it via a notification mechanism to the user (e.g., email or in-app notification).
[1646] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1647] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1648] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1649] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1650] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1651] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1652] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1653] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1654] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1655] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1656] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1657] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1658] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1659] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1660] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1661] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1662] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1663] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1664] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1665] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1666] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1667] The following is further disclosed regarding the above embodiment.
[1668] (Claim 1)
[1669] How to upload video files,
[1670] a means for inputting text data specifying a particular scene;
[1671] A means for analyzing a video file based on input text data and detecting specific scenes using a generative AI model;
[1672] means for extracting the detected scenes and generating a highlight video;
[1673] means for providing the generated highlight video to a user;
[1674] A system including:
[1675] (Claim 2)
[1676] 2. The system of claim 1, wherein the generative AI model includes a natural language processing model and an image and video analysis model.
[1677] (Claim 3)
[1678] 10. The system of claim 1, wherein the generative AI model uses facial recognition and / or behavior recognition algorithms to detect specific scenes.
[1679] "Example 1"
[1680] (Claim 1)
[1681] A means for uploading video data;
[1682] a means for inputting text data specifying a particular scene;
[1683] A means for analyzing video data based on input text data using a generative AI model to detect specific scenes;
[1684] a means for extracting the detected scenes and generating a highlight video;
[1685] a means for providing the generated highlight video to a user;
[1686] A system including:
[1687] (Claim 2)
[1688] 2. The system of claim 1, wherein the generative AI model includes a natural language processing model and an image and video analysis model.
[1689] (Claim 3)
[1690] 10. The system of claim 1, wherein the generative AI model uses person recognition and / or action recognition algorithms to detect specific scenes.
[1691] "Application Example 1"
[1692] (Claim 1)
[1693] How to upload video files,
[1694] a means for inputting text data specifying a particular scene;
[1695] A means for analyzing a video file based on input text data and detecting specific scenes using a generative AI model;
[1696] means for extracting the detected scenes and generating a highlight video;
[1697] a means for using cloud storage to provide the generated highlight video to a user;
[1698] a means for providing the generated highlight video by streaming;
[1699] A system including:
[1700] (Claim 2)
[1701] 2. The system of claim 1, wherein the generative AI model includes a natural language processing model and an image and video analysis model.
[1702] (Claim 3)
[1703] 10. The system of claim 1, wherein the generative AI model uses facial recognition and / or behavior recognition algorithms to detect specific scenes.
[1704] "Example 2: Combining Emotion Engines"
[1705] (Claim 1)
[1706] How to upload video files,
[1707] a means for inputting text data specifying a particular scene;
[1708] A means for analyzing a video file based on input text data and detecting specific scenes using a generative AI model;
[1709] means for collecting user emotion data;
[1710] means for analyzing emotion data and selecting scenes based on the user's emotional state;
[1711] means for extracting the detected scenes and generating a highlight video;
[1712] means for providing the generated highlight video to a user;
[1713] A system including:
[1714] (Claim 2)
[1715] 2. The system of claim 1, wherein the generative AI model includes a natural language processing model and an image and video analysis model.
[1716] (Claim 3)
[1717] 10. The system of claim 1, wherein the generative AI model uses facial recognition and / or behavior recognition algorithms to detect specific scenes.
[1718] "Application example 2 when combining emotion engines"
[1719] (Claim 1)
[1720] How to upload video files,
[1721] a means for inputting text data specifying a particular scene;
[1722] A means for analyzing a video file based on input text data and detecting specific scenes using a generative AI model;
[1723] means for extracting the detected scenes and generating a highlight video;
[1724] means for collecting user emotion data;
[1725] means for determining the importance of a scene based on emotion data;
[1726] means for providing the generated highlight video to a user;
[1727] A system including:
[1728] (Claim 2)
[1729] 2. The system of claim 1, wherein the generative AI model includes a natural language processing model and an image / video analysis model, and further includes an emotion engine that analyzes user emotion data.
[1730] (Claim 3)
[1731] 10. The system of claim 1, wherein the generative AI model utilizes facial recognition and behavior recognition algorithms and the user's emotional state to detect specific scenes. [Explanation of symbols]
[1732] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. How to upload video files, a means for inputting text data specifying a particular scene; A means for analyzing a video file based on input text data and detecting specific scenes using a generative AI model; means for extracting the detected scenes and generating a highlight video; means for providing the generated highlight video to a user; A system including:
2. 2. The system of claim 1, wherein the generative AI model includes a natural language processing model and an image / video analysis model.
3. 10. The system of claim 1, wherein the generative AI model uses facial recognition and / or behavior recognition algorithms to detect specific scenes.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A